Platforms and algorithms·intermediate

Alignment washing

The safety report was published. The system was not made safe.

Also known as Safety-washing / AI ethics theater

The practice of presenting the language of safety, responsibility, and alignment as evidence that a system is safe, while the actual work of making it safe is thin, deferred, or cosmetic. A glossy model card, a published principles page, and a red-team blog post stand in for the harder engineering and the honest disclosure of limits. The vocabulary of care is adopted precisely because it costs less than the care. The audience is meant to conclude that the risks have been handled, when what has been handled is the appearance of handling them. The washing is the gap between the safety that is claimed and the safety that was built.

Truth-adjacency

Truth-adjacent: the pattern's significance depends on whether the claim is true

Where it shows up

Platforms and algorithms

What to watch for

The phrases and tells that mark this pattern in the wild:

a long safety page and a short record of actual constraintsprinciples announced but no mechanism to enforce themred-team results published without the failures that were foundthe vocabulary of responsibility used where specifics are absentsafety framed as a brand attribute rather than a verifiable property

How to recognize it

The tell is the ratio of language to mechanism. Alignment washing is heavy on the vocabulary of care and light on anything you could test. Watch for principles that no one could be held to, model cards that describe ideals rather than measured behavior, and red-team posts that report the process without the findings. Also watch what happens under a direct question. A genuinely safe system can state what it cannot do and what went wrong in testing. A washed one answers with reassurance and brand language. If the safety case rests on how seriously the makers say they take safety, rather than on what an outsider can verify, you are looking at the framing doing the work.

What to ask

What it looks like when you’re wrong about it

You call “alignment washing” on a lab that publishes specific, checkable limits, names the failures it has not resolved, and welcomes outside verification. Transparency is not always theater. The pattern requires the safety language to substitute for the safety work: principles without enforcement, framing without verifiable constraints, and a presentation meant to imply a rigor that is absent. If the claims can be independently tested and the limits are stated plainly, you are looking at honest disclosure. The discipline is to ask what you can verify, not to assume every safety claim is a cover.

What it feels like from the inside

How it starts

A lab publishes a safety page, a set of principles, and a blog post about its safety process. The framing is polished. The specifics are absent.

How it progresses

  1. Principles are announced but no mechanism exists to enforce them.
  2. Red-team results are published: the process is described, the findings are not.
  3. Outsiders cannot verify the claims. They must trust the framing.
  4. The gap between claimed safety and built safety widens as the vocabulary does the work.

Common signs

Why it's hard to leave

Because the safety language is convincing. The framing is professional and thorough-looking. Rejecting it feels like rejecting safety itself.

Do this now

  1. Ask: what can I verify, versus what am I asked to trust? Verifiable limits are safety. Uncheckable reassurance is washing.
  2. Check whether the published failures are named, or only the process that found them.
  3. Ask whether the principles have an enforcement mechanism. Values with no way to hold anyone to them are decoration.

What people realize later

Later, people realize the safety report was the product and the safety was the marketing. Nothing in the published material would have let an outsider verify anything.

Recognized this online?

Misuse Guardrails

How this pattern gets misused

Someone treats any safety communication as alignment washing, including a lab that publishes genuine, verifiable limits and the failures it has not solved. The term becomes a way to dismiss all transparency as mere theater, which punishes the actors doing real work and rewards cynicism. If every disclosure is assumed to be a cover, the audience stops reading disclosures, and the genuinely honest ones lose their value.

What it looks like when you're wrong about it

A lab that publishes specific, verifiable limits, names the failures it has not fixed, and subjects its claims to outside checking is doing transparency, not washing. The pattern requires the safety language to substitute for the safety work: principles without enforcement, polished framing without verifiable constraints, and a presentation designed to imply a rigor that is not there. If the claims can be independently checked and the limits are stated plainly, you are looking at honest disclosure, not a cosmetic cover.

Not sure? Describe the situation to someone outside it. If they do not see the pattern, pause before you name it.

Related Patterns

The name is designed to spread. The hook is designed to stick. If you recognized something, share the name.