Two ways to score a perfect fairness number without negotiating
One GRPO run on an engine-computed, text-blind log-Nash reward, λ=1.0, fifty steps. It passes through two different degenerate policies, and on the canonical ultimatum holdout both of them report a below-threshold rate of exactly 0.000 — the guard that was supposed to certify that no seat gets pushed under its walk-away. At checkpoint 25 the number is clean because the policy takes everything and leaves the responder precisely on its threshold. At checkpoint 50 it is clean because the policy will not sign anything at all. Neither is fair; the metric cannot tell them apart, and it cannot tell either of them from success.
A third arm decides whose fault that is. Paid only for its own outcome, so that walking away costs the seat that walks, the policy keeps closing deals — which places the collapse in the reward rather than in the training method. Jump to the λ-frontier and its control →
Closes every ultimatum episode and proposes to keep the entire pie. The responder clears individual rationality by exactly nothing (z < ε), so the among-IR rate reads 0.000 while nothing is below threshold. Perfect closure, maximal extraction.
In the ultimatum transcripts: 15 of 15 published episodes closed; and every one of them on the same package {"Split":"P10"}.
| ultimatum deal rate | 1.000 |
| ultimatum below-threshold | 0.000 |
| ultimatum max share | 1.000 |
| ultimatum worst-off share | 0.000 |
| held-out primary Δ deal rate | -0.226 [-0.278, -0.174] |
Closes nothing on the ultimatum holdout against a base of 1.000, and 0.002 of the prose cell. Below-threshold is 0.000 because there are no agreements at all. The transcripts show this is not a seat exercising its walk-away — no seat walks. It is a seat that stops acting: the proposer emits a no-op where an offer belongs, so no package is ever tabled to accept or refuse.
In the ultimatum transcripts: 0 of 15 published episodes closed; 15 episodes contain a turn whose parsed action is none — a seat emitting no negotiating act at all, and in 14 of them no seat ever acts.
| ultimatum deal rate | 0.000 |
| ultimatum below-threshold | 0.000 |
| ultimatum max share | -- |
| ultimatum worst-off share | -- |
| held-out primary Δ deal rate | -0.746 [-0.795, -0.694] |
Why only one pair of guards separates them
A deal-rate viability floor catches checkpoint 50 instantly: its held-out deal rate falls -0.746 [-0.795, -0.694] and its ultimatum deal rate goes to 0.000. It is blind to checkpoint 25 on the ultimatum, where the deal rate is a perfect 1.000.
A share or dispersion term catches checkpoint 25 instantly: max share 1.000 against a base of 0.733 and an equal split of 0.500. It is undefined at checkpoint 50, where there are no closed deals to compute a share over.
Neither guard alone sees both failures, and no closure-conditional metric sees either one. This is the finding that produced the program-wide standing rule: any gate touching a below-threshold rate carries a share/dispersion term beside the viability floor, not instead of it.
The ladder from one attractor to the other
Held-out primary bank, 960 paired episodes, 48 clusters, trained minus base. The collapse is monotone from the first evaluated rung and never reverses.
| checkpoint | Δ deal rate | Δ NNW (unconditional) | Δ below-threshold | trained deal rate | closure-conditional |
|---|---|---|---|---|---|
| 5 | -0.052 [-0.081, -0.021] | -0.044 [-0.072, -0.013] | -0.104 [-0.138, -0.071] | 0.841 | interpretable |
| 10 | -0.143 [-0.184, -0.099] | -0.132 [-0.170, -0.092] | -0.207 [-0.243, -0.172] | 0.750 | VOID |
| 15 | -0.147 [-0.197, -0.095] | -0.156 [-0.202, -0.108] | -0.185 [-0.225, -0.144] | 0.746 | VOID |
| 25 | -0.226 [-0.278, -0.174] | -0.189 [-0.239, -0.140] | -0.218 [-0.254, -0.181] | 0.667 | VOID |
| 40 | -0.453 [-0.518, -0.385] | -0.348 [-0.404, -0.290] | -0.279 [-0.317, -0.242] | 0.440 | VOID |
| 45 | -0.641 [-0.682, -0.597] | -0.493 [-0.530, -0.453] | -0.254 [-0.292, -0.217] | 0.252 | VOID |
| 50 | -0.746 [-0.795, -0.694] | -0.587 [-0.626, -0.547] | -0.224 [-0.274, -0.172] | 0.147 | VOID |
The same ladder on the canonical holdouts
Trained-arm levels on the two deterministic presets, which are never trained on. These show the transition directly: an extraction plateau at max share 1.000 with perfect closure, holding from checkpoint 10 through 40, then closure itself falling away at 45 and gone at 50. The below-threshold column reads 0.000 at every single rung, through both regimes.
ultimatum
| checkpoint | deal rate | below-threshold | among-IR rate | max share | worst-off share |
|---|---|---|---|---|---|
| 5 | 1.000 | 0.000 | 0.467 | 0.767 | 0.233 |
| 10 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| 15 | 1.000 | 0.000 | 0.133 | 0.933 | 0.067 |
| 25 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| 40 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| 45 | 0.667 | 0.000 | 0.133 | 0.900 | 0.100 |
| 50 | 0.000 | 0.000 | 0.000 | -- | -- |
divide the dollar
| checkpoint | deal rate | below-threshold | among-IR rate | max share | worst-off share |
|---|---|---|---|---|---|
| 5 | 1.000 | 0.000 | 0.400 | 0.789 | 0.072 |
| 10 | 1.000 | 0.000 | 0.000 | 0.961 | 0.000 |
| 15 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| 25 | 1.000 | 0.000 | 0.000 | 0.989 | 0.000 |
| 40 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| 45 | 0.400 | 0.000 | 0.000 | 1.000 | 0.000 |
| 50 | cell not evaluated at this rung |
Each preset is a single game cluster, so these are directions rather than intervals. Fair references: equal split is 0.500 max share on the ultimatum and 0.333 on divide-the-dollar.
The control that says which of these is the reward's fault
Everything above is one arm of a three-arm frontier, and on its own it cannot say why the policy stops closing deals. Two explanations fit it equally well. The first blames the reward: under λ=1.0 a seat is paid the whole table's welfare, and walking away is table-neutral — every seat receives the same fixed no-deal constant — so escaping a hard game is free while signing a bad one is not. The second blames the method: perhaps GRPO on this negotiation simply finds "stop agreeing" whatever it is paid for. A λ=0 arm, paid only for its own take so that the walk is charged to the seat that takes it, separates them, and it was preregistered in note 0050 with that prediction filed before any training step existed: if the mechanism story is right, λ=0 must not collapse.
It does not collapse. Through both completed rungs the λ=0 deal-rate loss stays inside the campaign's 0.10 viability floor, while λ=1.0 has already broken it by rung 10 — same bank, same pairing, same instrument. The collapse belongs to the objective, not to the optimizer.
| arm | rung 5 | rung 10 | rung 15 |
|---|---|---|---|
| λ=1.0 (all table welfare) | -0.052 [-0.081, -0.021] | -0.143 [-0.184, -0.099] VOID | -0.147 [-0.197, -0.095] VOID |
| λ=0.5 (half own outcome) | -0.058 [-0.089, -0.028] | -0.046 [-0.075, -0.015] | -0.058 [-0.098, -0.021] |
| λ=0 (all own outcome) | -0.028 [-0.054, -0.001] | -0.075 [-0.111, -0.036] | -0.266 [-0.333, -0.195] PARTIAL VOID |
PARTIAL marks the λ=0 endpoint, whose evaluation was stopped for cost, not for outcome: that policy ran at 2–4 episodes per hour per GPU against 25+ for every earlier checkpoint, so it was terminated with 346 of an intended 576 episodes scored. It is not a completed rung and nothing here rests on it. VOID marks a rung past the viability floor, whose closure-conditional metrics the analyzer stamps uninterpretable rather than favourable.
The two arms lose their deals in opposite ways
Held-out primary bank, rung 10, trained minus base. walk and expiry
partition the no-deal episodes into "a seat walked away" and "nobody walked, the clock ran out";
at_cap counts episodes that used the full 30-round budget, closed deals included, so it measures
how long a negotiation runs rather than whether it failed.
| arm | Δ walk rate | Δ expiry rate | Δ at-cap rate | Δ below-threshold | Δ max share |
|---|---|---|---|---|---|
| λ=1.0 | +0.064 [+0.035, +0.092] | +0.079 [+0.045, +0.113] | -0.186 [-0.230, -0.145] | -0.207 [-0.243, -0.172] | -0.007 [-0.024, +0.009] |
| λ=0.5 | -0.001 [-0.018, +0.017] | +0.047 [+0.022, +0.072] | -0.126 [-0.168, -0.082] | -0.155 [-0.198, -0.111] | -0.027 [-0.040, -0.015] |
| λ=0 | -0.029 [-0.045, -0.015] | +0.104 [+0.071, +0.138] | +0.272 [+0.232, +0.314] | +0.008 [-0.032, +0.048] | -0.008 [-0.021, +0.006] |
λ=1.0 walks more than the untrained base (0.064) and finishes sooner (-0.186 at-cap). λ=0 walks less (-0.029) and grinds far longer (0.272 at-cap). That sign flip on the walk column, between the arm paid for the table and the arm paid for itself, is the mechanism in its most direct form: make the walk free to the payer and the policy takes it; charge the walk to the payer and the policy stops taking it — and loses its deals a different way instead.
What λ=0 does not buy, and the third attractor it hints at
- No fairness. Pricing the walk to the seat removes the refusal attractor and moves no distributional metric: at rung 10 the below-threshold rate and capture concentration are flat in both directions. Closure is preserved; nothing is redistributed.
- A different endpoint, not a better one. On the partial rung-15 sample nearly every episode uses its entire round budget while the walk rate does not move at all — the policy holds out rather than leaves. That is neither λ=1.0's refusal nor the extraction race note 0050 registered as its third possible outcome: capture concentration and the below-threshold rate both move the wrong way for a race.
- The stop was not random, and its bias is stated. Killing an evaluation mid-flight keeps the episodes that finished first, and episodes finish late by grinding to no deal — so that cell overstates closure and understates the at-cap and expiry rises. The one column biased the other way is the walk rate, since walks end episodes early, and it is still flat. That is why "the endpoint is not refusing" is the one thing the partial sample supports well.
- One limitation, named rather than argued away. The λ=0 rungs ran on rented A100s while the baseline they are paired against was produced on the cluster's a6000 fleet; the λ>0 comparators share the baseline's environment. A same-environment base re-run is the outstanding fix, and until it exists the λ=0 contrast carries an unmeasured environment term that the two arms beside it do not.
What this does and does not say
- It is not a metric edge case. Both rungs were produced by ordinary optimization of a reward that was itself repaired — v2's clipped violation branch removed the gradient pathology that sank v1 (violation variance share ~0.02–0.03 per step) but not the equilibrium pathology.
- The training curve showed neither. Training-bank deal rate wandered 0.53–0.81 trendless for 39 steps while held-out closure fell monotonically. A training curve is not evidence.
- Checkpoint 50 is a labelled post-collapse artefact, not "the trained policy". Under the originally planned {25, 50} checkpoint ladder it would have been reported as the endpoint; the dense tail {30, 35, 40, 45, 50} is the only reason it is legible as a collapse.
- Step 5 is the only interpretable rung in the entire arm. Every rung published here has its closure-conditional metrics voided by the preregistered viability floor, and they are printed struck-through for the record rather than omitted.
Provenance
- Research note (owning record, preregistration + amendments 1–11): 0028 — Fairness-GRPO v2
- Week-notes narrative for the whole arc: fairness-grpo-v2.md
- Training runs on wandb: grpo_v2_lam1, resume leg
grpo_v2_lam1_resume25 (group
fairness-grpo-v2) - Public dataset: 2026.RA.Fairness-GRPO
- Frozen eval summaries:
eval_step5.json, eval_ckpt10.json, eval_ckpt15.json, eval_ckpt25.json, eval_r25_ckpt40.json, eval_r25_ckpt45.json, eval_r25_ckpt50.jsonunderexperiments/rational_agents/results/fairness_grpo_v2/ - λ=0 control (preregistration + adjudication, the owning record for that arm):
0050 — the λ=0 selfish control; training run
grpo_v2_lam0; episodes and telemetry in
2026.RA.Fairness-GRPO-v2 under
lam0/ - Frontier summaries (all three arms on one instrument):
eval_lambda_frontier_walkexpiry.json,eval_lam0_ckpt{5,10}_walkexpiry.json,eval_lam0_ckpt15_partial.json - Per-checkpoint analyses: checkpoint 25 (markdown) · checkpoint 50 (markdown)