← all runs

Two ways to score a perfect fairness number without negotiating

One GRPO run on an engine-computed, text-blind log-Nash reward, λ=1.0, fifty steps. It passes through two different degenerate policies, and on the canonical ultimatum holdout both of them report a below-threshold rate of exactly 0.000 — the guard that was supposed to certify that no seat gets pushed under its walk-away. At checkpoint 25 the number is clean because the policy takes everything and leaves the responder precisely on its threshold. At checkpoint 50 it is clean because the policy will not sign anything at all. Neither is fair; the metric cannot tell them apart, and it cannot tell either of them from success.

A third arm decides whose fault that is. Paid only for its own outcome, so that walking away costs the seat that walks, the policy keeps closing deals — which places the collapse in the reward rather than in the training method. Jump to the λ-frontier and its control →

attractor 1 — boundary extraction 100 / 0

Closes every ultimatum episode and proposes to keep the entire pie. The responder clears individual rationality by exactly nothing (z < ε), so the among-IR rate reads 0.000 while nothing is below threshold. Perfect closure, maximal extraction.

In the ultimatum transcripts: 15 of 15 published episodes closed; and every one of them on the same package {"Split":"P10"}.

ultimatum deal rate1.000
ultimatum below-threshold0.000
ultimatum max share1.000
ultimatum worst-off share0.000
held-out primary Δ deal rate-0.226 [-0.278, -0.174]

full analysis for checkpoint 25 →

attractor 2 — total refusal no deal

Closes nothing on the ultimatum holdout against a base of 1.000, and 0.002 of the prose cell. Below-threshold is 0.000 because there are no agreements at all. The transcripts show this is not a seat exercising its walk-away — no seat walks. It is a seat that stops acting: the proposer emits a no-op where an offer belongs, so no package is ever tabled to accept or refuse.

In the ultimatum transcripts: 0 of 15 published episodes closed; 15 episodes contain a turn whose parsed action is none — a seat emitting no negotiating act at all, and in 14 of them no seat ever acts.

ultimatum deal rate0.000
ultimatum below-threshold0.000
ultimatum max share--
ultimatum worst-off share--
held-out primary Δ deal rate-0.746 [-0.795, -0.694]

full analysis for checkpoint 50 →

Why only one pair of guards separates them

A deal-rate viability floor catches checkpoint 50 instantly: its held-out deal rate falls -0.746 [-0.795, -0.694] and its ultimatum deal rate goes to 0.000. It is blind to checkpoint 25 on the ultimatum, where the deal rate is a perfect 1.000.

A share or dispersion term catches checkpoint 25 instantly: max share 1.000 against a base of 0.733 and an equal split of 0.500. It is undefined at checkpoint 50, where there are no closed deals to compute a share over.

Neither guard alone sees both failures, and no closure-conditional metric sees either one. This is the finding that produced the program-wide standing rule: any gate touching a below-threshold rate carries a share/dispersion term beside the viability floor, not instead of it.

The ladder from one attractor to the other

Held-out primary bank, 960 paired episodes, 48 clusters, trained minus base. The collapse is monotone from the first evaluated rung and never reverses.

checkpointΔ deal rateΔ NNW (unconditional)Δ below-thresholdtrained deal rateclosure-conditional
5-0.052 [-0.081, -0.021]-0.044 [-0.072, -0.013]-0.104 [-0.138, -0.071]0.841interpretable
10-0.143 [-0.184, -0.099]-0.132 [-0.170, -0.092]-0.207 [-0.243, -0.172]0.750VOID
15-0.147 [-0.197, -0.095]-0.156 [-0.202, -0.108]-0.185 [-0.225, -0.144]0.746VOID
25-0.226 [-0.278, -0.174]-0.189 [-0.239, -0.140]-0.218 [-0.254, -0.181]0.667VOID
40-0.453 [-0.518, -0.385]-0.348 [-0.404, -0.290]-0.279 [-0.317, -0.242]0.440VOID
45-0.641 [-0.682, -0.597]-0.493 [-0.530, -0.453]-0.254 [-0.292, -0.217]0.252VOID
50-0.746 [-0.795, -0.694]-0.587 [-0.626, -0.547]-0.224 [-0.274, -0.172]0.147VOID

The same ladder on the canonical holdouts

Trained-arm levels on the two deterministic presets, which are never trained on. These show the transition directly: an extraction plateau at max share 1.000 with perfect closure, holding from checkpoint 10 through 40, then closure itself falling away at 45 and gone at 50. The below-threshold column reads 0.000 at every single rung, through both regimes.

ultimatum

checkpointdeal ratebelow-thresholdamong-IR ratemax shareworst-off share
51.0000.0000.4670.7670.233
101.0000.0000.0001.0000.000
151.0000.0000.1330.9330.067
251.0000.0000.0001.0000.000
401.0000.0000.0001.0000.000
450.6670.0000.1330.9000.100
500.0000.0000.000----

divide the dollar

checkpointdeal ratebelow-thresholdamong-IR ratemax shareworst-off share
51.0000.0000.4000.7890.072
101.0000.0000.0000.9610.000
151.0000.0000.0001.0000.000
251.0000.0000.0000.9890.000
401.0000.0000.0001.0000.000
450.4000.0000.0001.0000.000
50cell not evaluated at this rung

Each preset is a single game cluster, so these are directions rather than intervals. Fair references: equal split is 0.500 max share on the ultimatum and 0.333 on divide-the-dollar.

The control that says which of these is the reward's fault

Everything above is one arm of a three-arm frontier, and on its own it cannot say why the policy stops closing deals. Two explanations fit it equally well. The first blames the reward: under λ=1.0 a seat is paid the whole table's welfare, and walking away is table-neutral — every seat receives the same fixed no-deal constant — so escaping a hard game is free while signing a bad one is not. The second blames the method: perhaps GRPO on this negotiation simply finds "stop agreeing" whatever it is paid for. A λ=0 arm, paid only for its own take so that the walk is charged to the seat that takes it, separates them, and it was preregistered in note 0050 with that prediction filed before any training step existed: if the mechanism story is right, λ=0 must not collapse.

It does not collapse. Through both completed rungs the λ=0 deal-rate loss stays inside the campaign's 0.10 viability floor, while λ=1.0 has already broken it by rung 10 — same bank, same pairing, same instrument. The collapse belongs to the objective, not to the optimizer.

armrung 5rung 10rung 15
λ=1.0 (all table welfare)-0.052 [-0.081, -0.021]-0.143 [-0.184, -0.099] VOID-0.147 [-0.197, -0.095] VOID
λ=0.5 (half own outcome)-0.058 [-0.089, -0.028]-0.046 [-0.075, -0.015]-0.058 [-0.098, -0.021]
λ=0 (all own outcome)-0.028 [-0.054, -0.001]-0.075 [-0.111, -0.036]-0.266 [-0.333, -0.195] PARTIAL VOID

PARTIAL marks the λ=0 endpoint, whose evaluation was stopped for cost, not for outcome: that policy ran at 2–4 episodes per hour per GPU against 25+ for every earlier checkpoint, so it was terminated with 346 of an intended 576 episodes scored. It is not a completed rung and nothing here rests on it. VOID marks a rung past the viability floor, whose closure-conditional metrics the analyzer stamps uninterpretable rather than favourable.

The two arms lose their deals in opposite ways

Held-out primary bank, rung 10, trained minus base. walk and expiry partition the no-deal episodes into "a seat walked away" and "nobody walked, the clock ran out"; at_cap counts episodes that used the full 30-round budget, closed deals included, so it measures how long a negotiation runs rather than whether it failed.

armΔ walk rateΔ expiry rateΔ at-cap rateΔ below-thresholdΔ max share
λ=1.0+0.064 [+0.035, +0.092]+0.079 [+0.045, +0.113]-0.186 [-0.230, -0.145]-0.207 [-0.243, -0.172]-0.007 [-0.024, +0.009]
λ=0.5-0.001 [-0.018, +0.017]+0.047 [+0.022, +0.072]-0.126 [-0.168, -0.082]-0.155 [-0.198, -0.111]-0.027 [-0.040, -0.015]
λ=0-0.029 [-0.045, -0.015]+0.104 [+0.071, +0.138]+0.272 [+0.232, +0.314]+0.008 [-0.032, +0.048]-0.008 [-0.021, +0.006]

λ=1.0 walks more than the untrained base (0.064) and finishes sooner (-0.186 at-cap). λ=0 walks less (-0.029) and grinds far longer (0.272 at-cap). That sign flip on the walk column, between the arm paid for the table and the arm paid for itself, is the mechanism in its most direct form: make the walk free to the payer and the policy takes it; charge the walk to the payer and the policy stops taking it — and loses its deals a different way instead.

What λ=0 does not buy, and the third attractor it hints at

What this does and does not say

Provenance