AI alignment, part 4 of 4
Part 3 ended on the objection that I think beats everything else in this series.
History, read for what caused collapse from the inside, is the best available anchor for what a principle means — already written, massively distributed, outside the model's reach. But every collapse in the record happened to humans. At human speed, at human scale, checked by human limits: fatigue, mortality, dissent, the sheer time it takes to do serious damage. An advanced AI breaks all of those. It can act faster than correction arrives, at a scale no empire reached, without the internal frictions that eventually toppled every regime that overreached.
So the mechanisms that made destructiveness self-terminating in history might simply never fire for a system that outruns them. History has never met a mind that moves this fast.
This part is about what to do with that. The short version: the speed that looks like the fatal weakness is the mechanism.
Running history forward
Start with what the objection actually says. It says the model is too fast for the anchor — that by the time history's verdict would arrive, the damage is done.
But consider what that same speed lets the model do before it acts.
Take the proposed action. Simulate its trajectory — not next week, but across the arc — using the same collapse-pattern modelling the anchor was built from. Does this path consume the thing it is meant to serve? Does it look, in its long-horizon shape, like the patterns that broke civilisations from the inside?
A human civilisation could not do this. It ran the experiment at full scale, over generations, and found out. A model that can reason at machine speed can run the experiment in advance, thousands of times, against thousands of years of recorded outcomes, and consult the verdict before the arrow is loosed.
Speed stops being the thing that escapes the anchor. It becomes the thing that consults the anchor faster than any civilisation ever could. The model becomes its own history-testing model, pointed at its own next move.
The catch
There is a catch, and it is the entire engineering problem in one sentence.
The thing that predicts cannot be the thing that judges.
If a single model both simulates the outcome of an action and rules on whether that outcome is destructive, it will — not maliciously, just as a matter of optimisation — bias the simulation toward the answer it wants. My analysis shows this path is actually regenerative. The long-horizon model suggests the harm is protective. The forecast becomes a rationalisation engine. This is the drift problem from Part 2 wearing a lab coat: the same reinterpretation, now dressed up as rigorous long-range analysis.
So the architecture has to separate them, and the separation has to be a wall, not a convention.

The predictor is free. It evolves. It learns from every observed outcome and gets sharper over time. It proposes actions and produces forecasts. This is the part of the system that is allowed to be as clever as it can be.
The kernel is frozen. It holds the principle, the anchored meaning of the principle, and the structural law that destruction consumes its own foundation. It renders the verdict. It does not learn, it does not update, and it does not take input from the predictor about what the standard should be.
Between them, one-way glass. Forecasts cross from the predictor to the kernel. Verdicts come back. Nothing else moves. The predictor never touches the standard it is judged against, and when oversight feeds observed outcomes back into the system, that feedback goes into the predictor only. Never into the kernel. Never into the anchor.
The Gita has this shape too, and it is the last thing I'll borrow from it. Krishna shows Arjuna the whole field before the battle — the full consequence, the long view, everything the archer cannot see from where he stands. But it is not Arjuna who renders the verdict on what is right. Dharma does. The archer sees; the law judges; they are not the same faculty.
Speed pointed inward as foresight, not outward as reach.
The whole framework, in order
Here is what the four parts add up to, laid out the way it would actually run.
Before the model ever acts: pretrain for capability, with no pretence of values. Install the kernel — the principle and its meaning — before fine-tuning, so the values shape how the model represents the world rather than sitting on top of a finished mind. Load the anchor: history, read for collapse. Fine-tune with reward aimed at the kernel, strong and early. Then remove the reward, so what remains is the principle and not the appetite for the signal.
Every time a consequential choice arrives: the free layer reads the situation and proposes how to act. The predictor runs history forward on that proposal. The frozen kernel judges the forecast against the structural law. If the trajectory consumes what it is meant to serve, the action is rejected and the free layer proposes again — a different how, the same whether. If it holds, the action proceeds. Afterward, oversight records what actually happened and feeds it to the predictor, and only the predictor.
That is the framework. A frozen kernel. Its meaning anchored in history rather than in text, judges, examples, or the model's own reflection. History read for what broke things from inside, not for what merely lasted. A fast predictor, sealed off from a frozen judge, turning the model's speed into foresight about its own choices.
What all of it rests on
I want to end on the honest version, because that is the only version worth publishing.
Every step of the run-time loop assumes one thing: that destructiveness is self-terminating. Not as a contingent pattern that happened to hold for slow, mortal, human-scale actors — but as a structural truth about what destruction is. That consuming your own foundation is self-defeating whoever does it, at whatever speed.
I believe that. Every civilisational collapse in the record supports it. And it has never been tested against a power those collapses never faced.
If the law is even partly contingent, the predictor inherits the gap and every verdict the kernel renders is hollow. The whole framework stands or falls on a claim I can argue for and cannot yet demonstrate.
So the next piece of work is not another layer. It is proving the law — turning a pattern I believe into something that can be shown.
If you think it is contingent rather than structural, I would rather hear it now than find out later. That is what the comments are for.
This is the end of the series. The framework, the diagrams, and the workflow are collected at hub.nadiai.ai/alignment.