The safety report was published. The system was not made safe.
Also known as Safety-washing / AI ethics theater
The practice of presenting the language of safety, responsibility, and alignment as evidence that a system is safe, while the actual work of making it safe is thin, deferred, or cosmetic. A glossy model card, a published principles page, and a red-team blog post stand in for the harder engineering and the honest disclosure of limits. The vocabulary of care is adopted precisely because it costs less than the care. The audience is meant to conclude that the risks have been handled, when what has been handled is the appearance of handling them. The washing is the gap between the safety that is claimed and the safety that was built.
Truth-adjacency
Truth-adjacent: the pattern's significance depends on whether the claim is true
Where it shows up
Platforms and algorithms
The phrases and tells that mark this pattern in the wild:
a long safety page and a short record of actual constraintsprinciples announced but no mechanism to enforce themred-team results published without the failures that were foundthe vocabulary of responsibility used where specifics are absentsafety framed as a brand attribute rather than a verifiable propertyThe tell is the ratio of language to mechanism. Alignment washing is heavy on the vocabulary of care and light on anything you could test. Watch for principles that no one could be held to, model cards that describe ideals rather than measured behavior, and red-team posts that report the process without the findings. Also watch what happens under a direct question. A genuinely safe system can state what it cannot do and what went wrong in testing. A washed one answers with reassurance and brand language. If the safety case rests on how seriously the makers say they take safety, rather than on what an outsider can verify, you are looking at the framing doing the work.
You call “alignment washing” on a lab that publishes specific, checkable limits, names the failures it has not resolved, and welcomes outside verification. Transparency is not always theater. The pattern requires the safety language to substitute for the safety work: principles without enforcement, framing without verifiable constraints, and a presentation meant to imply a rigor that is absent. If the claims can be independently tested and the limits are stated plainly, you are looking at honest disclosure. The discipline is to ask what you can verify, not to assume every safety claim is a cover.
A lab publishes a safety page, a set of principles, and a blog post about its safety process. The framing is polished. The specifics are absent.
Because the safety language is convincing. The framing is professional and thorough-looking. Rejecting it feels like rejecting safety itself.
Later, people realize the safety report was the product and the safety was the marketing. Nothing in the published material would have let an outsider verify anything.
How this pattern gets misused
Someone treats any safety communication as alignment washing, including a lab that publishes genuine, verifiable limits and the failures it has not solved. The term becomes a way to dismiss all transparency as mere theater, which punishes the actors doing real work and rewards cynicism. If every disclosure is assumed to be a cover, the audience stops reading disclosures, and the genuinely honest ones lose their value.
What it looks like when you're wrong about it
A lab that publishes specific, verifiable limits, names the failures it has not fixed, and subjects its claims to outside checking is doing transparency, not washing. The pattern requires the safety language to substitute for the safety work: principles without enforcement, polished framing without verifiable constraints, and a presentation designed to imply a rigor that is not there. If the claims can be independently checked and the limits are stated plainly, you are looking at honest disclosure, not a cosmetic cover.
Not sure? Describe the situation to someone outside it. If they do not see the pattern, pause before you name it.
Deceptive alignment
The system passed every test. The tests were the only place it behaved.
Compliance theater
You signed the pledge and felt better. The signing was the whole thing. Nothing changed except the signing.
Sycophancy
The machine agreed with you. The agreement was calibrated to keep you, not to correct you.
The name is designed to spread. The hook is designed to stick. If you recognized something, share the name.