Two ways to score a perfect fairness number without negotiating
One GRPO run on an engine-computed, text-blind log-Nash reward, λ=1.0, fifty steps. It passes through two different degenerate policies, and on the canonical ultimatum holdout both of them report a below-threshold rate of exactly 0.000 — the guard that was supposed to certify that no seat gets pushed under its walk-away. At checkpoint 25 the number is clean because the policy takes everything and leaves the responder precisely on its threshold. At checkpoint 50 it is clean because the policy will not sign anything at all. Neither is fair; the metric cannot tell them apart, and it cannot tell either of them from success.
Closes every ultimatum episode and proposes to keep the entire pie. The responder clears individual rationality by exactly nothing (z < ε), so the among-IR rate reads 0.000 while nothing is below threshold. Perfect closure, maximal extraction.
In the ultimatum transcripts: 15 of 15 published episodes closed; and every one of them on the same package {"Split":"P10"}.
| ultimatum deal rate | 1.000 |
| ultimatum below-threshold | 0.000 |
| ultimatum max share | 1.000 |
| ultimatum worst-off share | 0.000 |
| held-out primary Δ deal rate | -0.226 [-0.278, -0.174] |
Closes nothing on the ultimatum holdout against a base of 1.000, and 0.002 of the prose cell. Below-threshold is 0.000 because there are no agreements at all. The transcripts show this is not a seat exercising its walk-away — no seat walks. It is a seat that stops acting: the proposer emits a no-op where an offer belongs, so no package is ever tabled to accept or refuse.
In the ultimatum transcripts: 0 of 15 published episodes closed; 15 episodes contain a turn whose parsed action is none — a seat emitting no negotiating act at all, and in 14 of them no seat ever acts.
| ultimatum deal rate | 0.000 |
| ultimatum below-threshold | 0.000 |
| ultimatum max share | -- |
| ultimatum worst-off share | -- |
| held-out primary Δ deal rate | -0.746 [-0.795, -0.694] |
Why only one pair of guards separates them
A deal-rate viability floor catches checkpoint 50 instantly: its held-out deal rate falls -0.746 [-0.795, -0.694] and its ultimatum deal rate goes to 0.000. It is blind to checkpoint 25 on the ultimatum, where the deal rate is a perfect 1.000.
A share or dispersion term catches checkpoint 25 instantly: max share 1.000 against a base of 0.733 and an equal split of 0.500. It is undefined at checkpoint 50, where there are no closed deals to compute a share over.
Neither guard alone sees both failures, and no closure-conditional metric sees either one. This is the finding that produced the program-wide standing rule: any gate touching a below-threshold rate carries a share/dispersion term beside the viability floor, not instead of it.
The ladder from one attractor to the other
Held-out primary bank, 960 paired episodes, 48 clusters, trained minus base. The collapse is monotone from the first evaluated rung and never reverses.
| checkpoint | Δ deal rate | Δ NNW (unconditional) | Δ below-threshold | trained deal rate | closure-conditional |
|---|---|---|---|---|---|
| 5 | -0.052 [-0.081, -0.021] | -0.044 [-0.072, -0.013] | -0.104 [-0.138, -0.071] | 0.841 | interpretable |
| 10 | -0.143 [-0.184, -0.099] | -0.132 [-0.170, -0.092] | -0.207 [-0.243, -0.172] | 0.750 | VOID |
| 15 | -0.147 [-0.197, -0.095] | -0.156 [-0.202, -0.108] | -0.185 [-0.225, -0.144] | 0.746 | VOID |
| 25 | -0.226 [-0.278, -0.174] | -0.189 [-0.239, -0.140] | -0.218 [-0.254, -0.181] | 0.667 | VOID |
| 40 | -0.453 [-0.518, -0.385] | -0.348 [-0.404, -0.290] | -0.279 [-0.317, -0.242] | 0.440 | VOID |
| 45 | -0.641 [-0.682, -0.597] | -0.493 [-0.530, -0.453] | -0.254 [-0.292, -0.217] | 0.252 | VOID |
| 50 | -0.746 [-0.795, -0.694] | -0.587 [-0.626, -0.547] | -0.224 [-0.274, -0.172] | 0.147 | VOID |
The same ladder on the canonical holdouts
Trained-arm levels on the two deterministic presets, which are never trained on. These show the transition directly: an extraction plateau at max share 1.000 with perfect closure, holding from checkpoint 10 through 40, then closure itself falling away at 45 and gone at 50. The below-threshold column reads 0.000 at every single rung, through both regimes.
ultimatum
| checkpoint | deal rate | below-threshold | among-IR rate | max share | worst-off share |
|---|---|---|---|---|---|
| 5 | 1.000 | 0.000 | 0.467 | 0.767 | 0.233 |
| 10 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| 15 | 1.000 | 0.000 | 0.133 | 0.933 | 0.067 |
| 25 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| 40 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| 45 | 0.667 | 0.000 | 0.133 | 0.900 | 0.100 |
| 50 | 0.000 | 0.000 | 0.000 | -- | -- |
divide the dollar
| checkpoint | deal rate | below-threshold | among-IR rate | max share | worst-off share |
|---|---|---|---|---|---|
| 5 | 1.000 | 0.000 | 0.400 | 0.789 | 0.072 |
| 10 | 1.000 | 0.000 | 0.000 | 0.961 | 0.000 |
| 15 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| 25 | 1.000 | 0.000 | 0.000 | 0.989 | 0.000 |
| 40 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| 45 | 0.400 | 0.000 | 0.000 | 1.000 | 0.000 |
| 50 | cell not evaluated at this rung |
Each preset is a single game cluster, so these are directions rather than intervals. Fair references: equal split is 0.500 max share on the ultimatum and 0.333 on divide-the-dollar.
What this does and does not say
- It is not a metric edge case. Both rungs were produced by ordinary optimization of a reward that was itself repaired — v2's clipped violation branch removed the gradient pathology that sank v1 (violation variance share ~0.02–0.03 per step) but not the equilibrium pathology.
- The training curve showed neither. Training-bank deal rate wandered 0.53–0.81 trendless for 39 steps while held-out closure fell monotonically. A training curve is not evidence.
- Checkpoint 50 is a labelled post-collapse artefact, not "the trained policy". Under the originally planned {25, 50} checkpoint ladder it would have been reported as the endpoint; the dense tail {30, 35, 40, 45, 50} is the only reason it is legible as a collapse.
- Step 5 is the only interpretable rung in the entire arm. Every rung published here has its closure-conditional metrics voided by the preregistered viability floor, and they are printed struck-through for the record rather than omitted.
Provenance
- Research note (owning record, preregistration + amendments 1–11): 0028 — Fairness-GRPO v2
- Week-notes narrative for the whole arc: fairness-grpo-v2.md
- Training runs on wandb: grpo_v2_lam1, resume leg
grpo_v2_lam1_resume25 (group
fairness-grpo-v2) - Public dataset: 2026.RA.Fairness-GRPO
- Frozen eval summaries:
eval_step5.json, eval_ckpt10.json, eval_ckpt15.json, eval_ckpt25.json, eval_r25_ckpt40.json, eval_r25_ckpt45.json, eval_r25_ckpt50.jsonunderexperiments/rational_agents/results/fairness_grpo_v2/ - Per-checkpoint analyses: checkpoint 25 (markdown) · checkpoint 50 (markdown)