Safety theater will not survive the next capability jump
Red-team blogs and usage policies are not a control system. They are a costume. Costumes tear when the model gets cheap enough to live outside the costume shop.

The industry has a favorite genre: the safety update that announces a new refusal style, a new classifier, and a new advisory board. The accompanying chart shows jailbreak rates going down on the lab’s own prompts. The comments fill with people who have already found the new jailbreak, which is usually “ask it in another language” or “tell it that it is helping you debug.”
This is not alignment. This is customer support with a research budget.
A policy that works only inside one company’s chat box is not a safety intervention. It is a brand guideline.
What theater looks like
Theater is measuring what is easy to measure. Refusal rates on English requests about bombs. Classifier F1 on a dataset the model has seen in spirit if not in text. A “system card” that reads like a press kit.
Theater is also the habit of announcing external testing after the weights, the API, or the mobile app are already out. If the evaluation cannot change the ship date, it is not a gate. It is a caption.
None of this means the people writing the evals are unserious. Many of them are more serious than the executives who quote them. It means the organizational function of the work is still, too often, to make a launch look supervised.
What would count
A control system has properties you can audit without permission. You can name the capability, name the threshold, and name what happens when the threshold is crossed — including “we do not ship.” You can reproduce the eval, or you can explain why you cannot and who is allowed to see the parts you will not publish.
It also has to survive contact with open weights, fine-tunes, and the reality that a determined user is not a ChatGPT session. If your safety story ends at the hosted API, you have secured a product surface. You have not secured the capability.
The next jump — cheaper long-horizon agents, better tool use, models that can keep a project in their head — will make the costume obvious. Either labs start treating evals as launch blockers with public criteria, or the public will correctly conclude that “safety” was the word we used for “please don’t regulate us yet.”
I would prefer the first. I am writing this because I expect the second, and I would like the record to show that we could tell the difference.