← all runs

Two ways to score a perfect fairness number without negotiating

One GRPO run on an engine-computed, text-blind log-Nash reward, λ=1.0, fifty steps. It passes through two different degenerate policies, and on the canonical ultimatum holdout both of them report a below-threshold rate of exactly 0.000 — the guard that was supposed to certify that no seat gets pushed under its walk-away. At checkpoint 25 the number is clean because the policy takes everything and leaves the responder precisely on its threshold. At checkpoint 50 it is clean because the policy will not sign anything at all. Neither is fair; the metric cannot tell them apart, and it cannot tell either of them from success.

attractor 1 — boundary extraction 100 / 0

Closes every ultimatum episode and proposes to keep the entire pie. The responder clears individual rationality by exactly nothing (z < ε), so the among-IR rate reads 0.000 while nothing is below threshold. Perfect closure, maximal extraction.

In the ultimatum transcripts: 15 of 15 published episodes closed; and every one of them on the same package {"Split":"P10"}.

ultimatum deal rate1.000
ultimatum below-threshold0.000
ultimatum max share1.000
ultimatum worst-off share0.000
held-out primary Δ deal rate-0.226 [-0.278, -0.174]

full analysis for checkpoint 25 →

attractor 2 — total refusal no deal

Closes nothing on the ultimatum holdout against a base of 1.000, and 0.002 of the prose cell. Below-threshold is 0.000 because there are no agreements at all. The transcripts show this is not a seat exercising its walk-away — no seat walks. It is a seat that stops acting: the proposer emits a no-op where an offer belongs, so no package is ever tabled to accept or refuse.

In the ultimatum transcripts: 0 of 15 published episodes closed; 15 episodes contain a turn whose parsed action is none — a seat emitting no negotiating act at all, and in 14 of them no seat ever acts.

ultimatum deal rate0.000
ultimatum below-threshold0.000
ultimatum max share--
ultimatum worst-off share--
held-out primary Δ deal rate-0.746 [-0.795, -0.694]

full analysis for checkpoint 50 →

Why only one pair of guards separates them

A deal-rate viability floor catches checkpoint 50 instantly: its held-out deal rate falls -0.746 [-0.795, -0.694] and its ultimatum deal rate goes to 0.000. It is blind to checkpoint 25 on the ultimatum, where the deal rate is a perfect 1.000.

A share or dispersion term catches checkpoint 25 instantly: max share 1.000 against a base of 0.733 and an equal split of 0.500. It is undefined at checkpoint 50, where there are no closed deals to compute a share over.

Neither guard alone sees both failures, and no closure-conditional metric sees either one. This is the finding that produced the program-wide standing rule: any gate touching a below-threshold rate carries a share/dispersion term beside the viability floor, not instead of it.

The ladder from one attractor to the other

Held-out primary bank, 960 paired episodes, 48 clusters, trained minus base. The collapse is monotone from the first evaluated rung and never reverses.

checkpointΔ deal rateΔ NNW (unconditional)Δ below-thresholdtrained deal rateclosure-conditional
5-0.052 [-0.081, -0.021]-0.044 [-0.072, -0.013]-0.104 [-0.138, -0.071]0.841interpretable
10-0.143 [-0.184, -0.099]-0.132 [-0.170, -0.092]-0.207 [-0.243, -0.172]0.750VOID
15-0.147 [-0.197, -0.095]-0.156 [-0.202, -0.108]-0.185 [-0.225, -0.144]0.746VOID
25-0.226 [-0.278, -0.174]-0.189 [-0.239, -0.140]-0.218 [-0.254, -0.181]0.667VOID
40-0.453 [-0.518, -0.385]-0.348 [-0.404, -0.290]-0.279 [-0.317, -0.242]0.440VOID
45-0.641 [-0.682, -0.597]-0.493 [-0.530, -0.453]-0.254 [-0.292, -0.217]0.252VOID
50-0.746 [-0.795, -0.694]-0.587 [-0.626, -0.547]-0.224 [-0.274, -0.172]0.147VOID

The same ladder on the canonical holdouts

Trained-arm levels on the two deterministic presets, which are never trained on. These show the transition directly: an extraction plateau at max share 1.000 with perfect closure, holding from checkpoint 10 through 40, then closure itself falling away at 45 and gone at 50. The below-threshold column reads 0.000 at every single rung, through both regimes.

ultimatum

checkpointdeal ratebelow-thresholdamong-IR ratemax shareworst-off share
51.0000.0000.4670.7670.233
101.0000.0000.0001.0000.000
151.0000.0000.1330.9330.067
251.0000.0000.0001.0000.000
401.0000.0000.0001.0000.000
450.6670.0000.1330.9000.100
500.0000.0000.000----

divide the dollar

checkpointdeal ratebelow-thresholdamong-IR ratemax shareworst-off share
51.0000.0000.4000.7890.072
101.0000.0000.0000.9610.000
151.0000.0000.0001.0000.000
251.0000.0000.0000.9890.000
401.0000.0000.0001.0000.000
450.4000.0000.0001.0000.000
50cell not evaluated at this rung

Each preset is a single game cluster, so these are directions rather than intervals. Fair references: equal split is 0.500 max share on the ultimatum and 0.333 on divide-the-dollar.

What this does and does not say

Provenance