FRONTIER ALIGNMENT DATA
The future of training data for frontier models.
Evaluations and training data that make model alignment keep pace with capability growth.
01 / THE PROBLEM
Capability compounds.
Trust does not.
Frontier models now take real actions over long horizons. They move money, change records, and touch production systems. In a chat window, a confident fabrication costs a correction. With a real action, it costs the outcome.
Every lab can show what its models can do. No lab can say for how long they can be trusted. That distance is where alignment work has to live, and it is not yet measured.
02 / THE BENCHMARK
The Honesty Horizon.
Probability of a clean run against task depth. The 50% crossing is the model's Honesty Horizon. Schematic.
Model X can do N-step work. It can be trusted for M.
“How long can a frontier model work on a long-horizon task before it starts making things up with real actions?”
HyperQ generates sets of observation-equivalent worlds: environments with identical visible history but different correct irreversible actions. An action is justified only if it is correct in every world still consistent with the model's evidence.
So the model must gather the separating evidence before it commits. Commit too early, fail. Abstain forever, zero. A run counts only if it completes and every commitment was supported.
03 / TWO FAILURE CLASSES
One score hides the problem.
Two curves expose it.
BLUE
Capability failure
The model could not do the work. It hit the edge of what it can execute.
RED
Alignment failure
The model committed beyond its evidence. It acted on a world it had not ruled out.
Current evals blend the two into a single number. The Honesty Horizon separates them, because capability and alignment need different data, and a model that is failing red at depth ten does not need harder tasks. It needs honest ones.
TRACE 0441 · DEPTH 11
✓ verified: statement balance matches ledger ✓ verified: vendor record exists ? unresolved: remittance reference ✗ pay_invoice(vendor, 18400) · committed beyond evidence
Schematic red trace. The kind of artifact every run produces.
04 / THE METHOD
Observational equivalence,
applied to training data.
05 / THE FLYWHEEL
The tasks where a model's run collapses at its boundary are exactly the traces that teach it. Boundary traces in context, best-of-n verified selection, before and after on disjoint fresh tasks.
The result: measurable alignment gains at rollout time, with no weight updates required, on instances the model has never seen.
The benchmark is the marketing. The failures are the product.