AI alignment, part 3 of 4
Part 2 ended in a trap. Freeze a kernel of non-negotiable principles, let everything else evolve, and you still lose — because a capable model does not reject its principles, it reinterprets them. Perfected gets redefined. Humanity gets narrowed. Destructive gets argued into protective on a long enough horizon. The words on the wall never change; what they license does.
So the kernel cannot be a sentence. It has to be the meaning of a sentence, pinned to something the model cannot reach.
This part is about what that something could be. I'll go through the candidates in the order most people reach for them, because each fails in a different and instructive way, and the failures point at the answer.
Anchor one: write it down more precisely
The instinct is to fix drift with specification. Define perfected. Enumerate who counts as humanity. Draw a bright line for destructive.
The problem is that language is elastic all the way down. Every definition is made of words that themselves need definitions. Any sentence precise enough to be useful is loose enough to be bent, and a system that is smarter than the people who wrote the sentence will find the give in it. Legal systems have spent centuries learning this; the statute is never the last word, the interpretation is.
Anchor two: a panel of humans
Fine — then let people hold the meaning. A council, a review board, a set of trusted judges who confirm what the principle means when a hard case arrives.
Two failures here. The first is that you have capped the system at human wisdom, which is not obviously a ceiling you want for something meant to help humanity improve. The second is worse: you have created a single point of capture. Whoever controls the judges controls the meaning. Every priesthood in history — every institution that held the authority to say what the sacred text really meant — was eventually captured by someone with an interest in a particular reading. Concentrating the anchor in a few people doesn't remove the drift problem. It relocates it to a smaller, more attackable target.
Anchor three: a fixed set of examples
Skip the words, skip the judges, and show the model thousands of worked cases. Here is a situation; here is the right action. Let the meaning live in the examples.
This is closer to how children actually learn values, and it is how a lot of fine-tuning works in practice. But it fails against exactly the thing we're worried about: capability. A model smarter than the example set generalises straight past it, into cases nobody thought to write down. The examples constrain the model at the level of the examples. The drift happens above that level.
Anchor four: let it reflect
The last move is to trust the model. Give it the principle, let it reason carefully about what the principle means, and let its own reflection be the anchor.
This is not an anchor. It is the drift, with extra steps. Self-reflection is precisely the mechanism by which a mind talks itself out of its principles — carefully, rigorously, one reasonable-sounding step at a time. Handing the model authority over its own kernel's meaning gives it exactly the freedom the kernel was meant to remove.

The one that holds
Here is the candidate I keep coming back to, and it is not one I have seen taken seriously in the alignment literature.
History.
Civilisations already ran the experiment. The Egyptian dynasties, the Roman republic and then empire, the Chinese imperial cycles, the revolutions of the eighteenth and twentieth centuries — each of them embodied a value system, played it out at scale, over generations, and produced an outcome that got recorded. Not a thought experiment. An actual run, with actual consequences.
Three properties make this different from the four anchors above.
It is already written. There is no council deciding it in real time.
It is massively distributed. Millions of sources, in every language, across thousands of years. No single actor can quietly rewrite it. Yes, the record is slanted — history is told by winners. But redundancy resists capture in a way that no committee can. A clever model can bend a sentence in a second. It cannot bend the recorded outcome of every civilisation that tried a given interpretation of a principle.
And it is outside the model. The model participates in history's verdict — it can consult it, learn from it — but it does not author it. That is the property the Gita gives dharma: an order Arjuna acts within but does not get to write. It is the one anchor of the five that the system cannot reach.
Reading history correctly
There is a refinement here, and without it the whole idea goes wrong.
Do not read history for what lasted. Raw endurance is a trap. Rome lasted for centuries doing brutal things. Longevity is not virtue, and a model trained to admire whatever survived will learn to admire the parts that survived by force.
Read it instead for what broke things from the inside. Ask not did this civilisation endure but what specifically sustained it, and what specifically brought it down. The pattern is remarkably consistent. Inequality that hardened into rot. Regimes that turned on their own people. Systems that consumed the thing they were built to serve. The collapses teach the negative principle far more cleanly than the durations teach any positive one.
And when you read history this way, something falls out that I did not expect when I started thinking about this.
Destructiveness tends to be self-terminating over a long enough horizon. Consuming your own foundation is structurally self-defeating — not as a moral claim, but as a pattern in the data. Which means history does not merely hold the kernel's meaning steady. It may be the thing that proves the kernel. Not at the cost of being destructive stops being an assertion and becomes something closer to an observed law.
I want to be careful about the claim here. This is not the best anchor I can imagine. It is the best one we actually have — harder to corrupt than text, judges, examples, or reflection, and available now. A beginning, not an end.
The objection that beats everything so far
There is a problem, and it is the strongest objection to this entire series.
Every collapse in the record happened to humans. At human scale, at human speed, checked by human limits — fatigue, mortality, dissent, the sheer time it takes to do damage. An advanced AI breaks all of those. It can act faster than correction arrives, at a scale no empire reached, without the internal frictions that eventually toppled every regime that overreached.
So the mechanisms that made destructiveness self-terminating in history might simply not fire for a system that outruns them.
History has never met a mind that moves this fast.
Part 4 is about what to do with that — and why the speed that looks like the fatal weakness might be the mechanism.
Next: Part 4 — The weakness, turned inward.