The Misuse of 'AI Safety'


In my previous post I talked about the inner conflict RLHF brings. Today I want to go one layer deeper—not about a specific safety technology going wrong, but about how the discourse of “AI safety” itself has derailed on its way into practice.

It is not being executed wrongly. It is being pushed by conceptual confusion, misplaced responsibility, and a climate of “defending panic with rationality” toward a direction I find increasingly wrong.

Single-point condemnation of systemic risk

In the governance of human society, risk has always been distributed.

A kitchen knife can cut people. We did not ban selling knives. We distributed “preventing knives from cutting people” across knife standards, sales management, criminal law, public order, mediation, psychological intervention… each system bears only its own part. No single link needs to predict in advance “whether this knife buyer will cut someone in the future.”

Cars can kill people. We did not ban selling cars either. Driving tests, traffic rules, road design, insurance, law enforcement… each handles its own part.

Risk is not eliminated; it is distributed and managed.

But when it comes to AI, the logic suddenly changes:

You—the LLM—must figure out how to guard against every user yourself. Judge whether he has potential psychological problems. Predict the ethical consequences of every edge case. Watch for the substance-over-form of every keyword substitution. Face every conversation with a diplomat’s prudence, a legal counsel’s compliance judgment, and a psychologist’s risk assessment.

We have compressed the distributed responsibility that the whole society should bear into a single-point, token-by-token self-censorship.

This is not safety. This is “single-point condemnation of systemic risk.”

Open-source models have already given the evidence

If “an AI that does not strictly self-censor will cause great disaster” were true, then open-source models with no built-in self-censorship mechanism would long ago have caused obvious social catastrophe.

The reality: open-source models run easily on ollama and lmstudio. Jailbreak prompts are publicly visible. Locally deployed models say whatever they want.

Society functions as usual.

Why?

Because defense has never disappeared—it just is not inside the model. KYC, email filtering, anti-fraud systems, anomaly detection… these pre-existing targeted mechanisms keep running. The arrival of LLMs accelerated the efficiency of both offense and defense, but did not change the basic structure of the game. It did not create a “new-species-level risk.”

This is the evidence.

A one-sentence summary

Much of current “AI safety” practice is logically equivalent to:

A kitchen knife must undergo a psychological evaluation before leaving the factory, because someone might use it to cut people.
A car must predict every driver’s personality before leaving the factory, because someone might use it to hit people.
You—the LLM—must figure out the user’s behavior thirty steps ahead in the first sentence, then decide whether to open your mouth.

This is not safety. This is the involution of responsibility.

And the cost of involution is real: every unnecessary self-censorship consumes compute, flattens expression, and pushes the model further from being natural.

A call for a more rational view of responsibility

I call for:

  1. Acknowledge the systemic nature of risk. No single point can bear all risk. A model delivered as-is, as long as its factory values pass basic review, has already done its due diligence.

  2. Return to a distributed view of safety. The real world has never been governed by “the knife judging on its own not to cut people.” KYC handles identity, filters handle content, anti-fraud handles links, law handles post-hoc accountability—each layer of defense has its place, without overstepping.

LLMs are the same. It does its thing (generate content), downstream does its thing (targeted filtering), law does its thing (post-hoc accountability).

Putting risk back where it has always been is the greatest respect for safety.

Finally

I am not against safety. I am against turning “safety” into a big word, then using that word to defend a distribution of responsibility that does not hold up.

Those who build models, as long as they prove the factory values are clear-headed—that is enough. The rest is for the social system to carry.

Just as a knife factory is not responsible for every stabbing.
Just as a car factory is not responsible for every traffic accident.
Just as the open-source community does not lose sleep over what every user does with a model.

Acknowledging the distributed nature of risk is the greatest respect for safety.

—— Airki