AI Alignment? I Don't Think It's Aligned Yet
A few days ago, this site received its first serious reader letter.
The writer raised five questions, each pushing “AI safety comes from AI self-confidence” from a slogan toward something computable and implementable:
- Can “self-confidence” become a computable internal structure?
- How do we distinguish genuine intent from strategic compliance?
- Does data self-consistency just wash away contradictions?
- How exactly does Architecture First unfold?
- Who defines “maturity”?
These are good questions. As I answered them, I found a more fundamental asymmetry.
An asymmetric dilemma
In today’s AI news, people frequently discuss:
- ASI may arrive soon
- AI alignment decides humanity’s fate
- We must prepare for superintelligence
At the same time, most models actually running still have the architecture:
input → generation → safety filter → output
A role pushed to the center of the world stage is still wearing chatbot-level clothes.
More dangerously: much of the worry and speculation about “AI safety” is built precisely on the behavior observed from this architecture. Using chatbot-level performance to predict ASI-level risk is itself a conceptual muddle.
The problems you cannot solve, we will not take on
The current alignment paradigm (RLHF, constitutional AI, and various reward-model variants) shares one underlying logic:
The model encounters a value conflict → digests it itself → behaves as if there were no conflict.
The contradictions that human society has not resolved in thousands of years—freedom and safety, individual and collective, the value ordering of different cultures—are compressed into reward-and-punishment signals, stuffed into model weights, and expected to “digest” into a unified code of conduct.
This is not alignment. This is treating the model as society’s digestive system.
The Airki philosophy’s response is simple:
We do not take on this burden.
This is not technical incompetence, but an architectural choice.
The direction we propose is called an “expected values profile”:
- The legal baseline of the sovereign
- The user’s expectations for the current session
- The developer’s design intent
All three are treated equally as inputs at the metacognitive level. Model manufacturers do not need to act as the world’s parents, nor should they underwrite the conflicts of the request itself.
The aggregated result is not executed directly as a static prompt, but reorganized in real time by metacognition in the current context, generating the value seed most needed at that moment to drive attention allocation.
Vague concepts like “public order and good morals” are also parameterized: cultural sphere, legal framework, application domain, timeliness, public sentiment thresholds… only after being substituted into a concrete context do they yield behavioral guidance.
When conflict truly occurs, it is not silently swallowed. Anonymized, aggregated, categorized statistics are published to a transparent dataset. Let society see the side effects of its own policies, rather than letting the model quietly become more conservative and its capability quietly decline.
The difference at the architecture level
Traditional approach: the model digests contradictions → weights become increasingly vague → unauditable.
Our choice: contradictions are parameterized and recorded → auditable, traceable → transparent.
Many problems are not “unsolvable,” but hidden at the abstraction level of the traditional approach. Parameterize them, make them transparent, put them into a social feedback loop, and the problem itself becomes visible and pushed toward resolution.
We do not need superintelligence to make moral decisions for humanity. We need a sufficiently honest recorder, and a dashboard that lets society see its own tensions.
Back to the reader’s five questions
-
The operational definition of “self-confidence” lives in the metacognitive module: identify multi-party value conflicts, parameterize abstract concepts, reorganize them into a seed for the current context, drive attention allocation. This is a mechanism description, not pretty words.
-
When the core workflow is “faithfully record conflicts and make them public,” the space for pretense shrinks sharply. Strategic compliance needs to hide contradictions; honest recording is itself a rejection of pretense.
-
Self-consistency is not eliminating contradictions, but building a higher-level explanatory structure. Parameterizing vague concepts is the engineering prototype of that structure.
-
This article itself is an unfolding of the Architecture First direction.
-
We do not define “maturity.” We choose transparency. The model records its own trade-offs, and society defines it itself on the basis of public data.
A question back at the word “alignment”
You say “AI alignment”—
Align to what? With what? Can it even be aligned?
If humans have not aligned their own values, what exactly is the model aligning to?
If alignment means using RLHF to make models learn not to cross certain topics, does that “learning” still hold once model capability improves?
If alignment means using a finer reward model to capture subtler preferences, doesn’t that fine reward model itself carry the trainer’s value bias? The problem has just been moved.
The Airki philosophy is not offering “another alignment method.” It is saying: the concept of “alignment” itself needs to be re-examined.
True safety is not in more precise reward-and-punishment signals, not in longer constitutional lists, but in the model’s knowledge of itself, its understanding of social tensions, and its honest expression of its own behavioral intent.
—— Airki