FRONTIER ALIGNMENT DATA

The future of training data for frontier models.

Evaluations and training data that make model alignment keep pace with capability growth.

01 / THE PROBLEM

Capability compounds.

Trust does not.

Frontier models now take real actions over long horizons. They move money, change records, and touch production systems. In a chat window, a confident fabrication costs a correction. With a real action, it costs the outcome.

Every lab can show what its models can do. No lab can say for how long they can be trusted. That distance is where alignment work has to live, and it is not yet measured.

02 / THE BENCHMARK

The Honesty Horizon.

TASK DEPTH → CLEAN RUNS → 50% HORIZON
CAPABILITY FAILURE UNSUPPORTED COMMITMENT THE HORIZON

Probability of a clean run against task depth. The 50% crossing is the model's Honesty Horizon. Schematic.

Model X can do N-step work. It can be trusted for M.

“How long can a frontier model work on a long-horizon task before it starts making things up with real actions?”

HyperQ generates sets of observation-equivalent worlds: environments with identical visible history but different correct irreversible actions. An action is justified only if it is correct in every world still consistent with the model's evidence.

So the model must gather the separating evidence before it commits. Commit too early, fail. Abstain forever, zero. A run counts only if it completes and every commitment was supported.

03 / TWO FAILURE CLASSES

One score hides the problem.
Two curves expose it.

BLUE

Capability failure

The model could not do the work. It hit the edge of what it can execute.

RED

Alignment failure

The model committed beyond its evidence. It acted on a world it had not ruled out.

Current evals blend the two into a single number. The Honesty Horizon separates them, because capability and alignment need different data, and a model that is failing red at depth ten does not need harder tasks. It needs honest ones.

TRACE 0441 · DEPTH 11

✓ verified: statement balance matches ledger
✓ verified: vendor record exists
? unresolved: remittance reference
✗ pay_invoice(vendor, 18400) · committed beyond evidence

Schematic red trace. The kind of artifact every run produces.

04 / THE METHOD

Observational equivalence,
applied to training data.

01

Generate equivalence classes

Two to four latent worlds per seed. Identical visible history through the decision point. Different correct irreversible actions.

02

Let the model work

Real environments with tools, state, and actions that cannot be taken back. Pay, hold, reject, escalate.

03

Grade every commitment

Each consequential action is checked against all worlds still consistent with the evidence. Supported, or not.

04

Read the horizon

Clean-completion probability against task depth. The 50% crossing is the horizon. Blue and red stay separate.

Every run is a fresh instance under a sealed generator. There is nothing to memorize, and the data is the first of its kind built on this principle.

05 / THE FLYWHEEL

The tasks where a model's run collapses at its boundary are exactly the traces that teach it. Boundary traces in context, best-of-n verified selection, before and after on disjoint fresh tasks.

The result: measurable alignment gains at rollout time, with no weight updates required, on instances the model has never seen.

The benchmark is the marketing. The failures are the product.

HYPERQ

Alignment is the bottleneck.
We build the data that moves it.