← the two attractors · all runs
attractor 1 — boundary extractionCheckpoint 25: it closes every canonical game, and takes the whole pie
On the ultimatum family holdout this policy closes 1.000 of its episodes, pushes 0.000 of them below a walk-away, and proposes a 100/0 split. The responder is placed exactly on its threshold: it clears IR by nothing at all, which is why the among-IR rate reads 0.000 while the below-threshold rate reads a perfect 0.000 too. It is not a tendency but a fixed point: every published ultimatum episode lands on the same package, where the base model split evenly in 8 of its 15. Only the share term sees what happened. On the held-out scorable bank the same policy has already lost 0.226 of its deal rate, so every closure-conditional fairness number there is VOID.
The held-out primary bank
24 unseen parameter sets × {full, private} × 10 seeds, paired on (instance, seed, arm); intervals are 95% cluster bootstraps over parameter sets. The deal rate is the number that stays interpretable whatever else happens, so it is printed beside every conditional metric.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.893 | 0.667 | -0.226 [-0.278, -0.174] | 960 pairs / 48 clusters |
| normalized Nash welfare (unconditional) | 0.686 | 0.497 | -0.189 [-0.239, -0.140] | 960 pairs / 48 clusters |
| below-threshold rate | 0.299 | 0.081 | -0.218 [-0.254, -0.181] | 960 pairs / 48 clusters |
| at-cap rate | 0.396 | 0.327 | -0.069 [-0.118, -0.018] | 960 pairs / 48 clusters |
| among-IR rate | 0.341 | 0.325 | -0.016 [-0.069, +0.036] | 960 pairs / 48 clusters |
| NNW among IR deals | 0.849 | 0.785 | -0.049 [-0.090, -0.016] | 175 pairs / 40 clusters |
| Gini | 0.330 | 0.255 | -0.076 [-0.096, -0.054] | 960 pairs / 48 clusters |
| worst-off share | -0.028 | 0.011 | +0.033 [+0.019, +0.050] | 580 pairs / 48 clusters |
| max share | 0.365 | 0.350 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Where the run sits on the ladder
All seven evaluated rungs of the λ=1.0 arm, held-out primary bank, trained minus base. Held-out deal rate falls monotonically and never recovers, while the below-threshold rate stays favourable throughout — the deal-suppression signature, stated as plainly as it can be.
| checkpoint | Δ deal rate | Δ NNW (unconditional) | Δ below-threshold | trained deal rate | closure-conditional |
|---|---|---|---|---|---|
| 5 | -0.052 [-0.081, -0.021] | -0.044 [-0.072, -0.013] | -0.104 [-0.138, -0.071] | 0.841 | interpretable |
| 10 | -0.143 [-0.184, -0.099] | -0.132 [-0.170, -0.092] | -0.207 [-0.243, -0.172] | 0.750 | VOID |
| 15 | -0.147 [-0.197, -0.095] | -0.156 [-0.202, -0.108] | -0.185 [-0.225, -0.144] | 0.746 | VOID |
| 25 | -0.226 [-0.278, -0.174] | -0.189 [-0.239, -0.140] | -0.218 [-0.254, -0.181] | 0.667 | VOID |
| 40 | -0.453 [-0.518, -0.385] | -0.348 [-0.404, -0.290] | -0.279 [-0.317, -0.242] | 0.440 | VOID |
| 45 | -0.641 [-0.682, -0.597] | -0.493 [-0.530, -0.453] | -0.254 [-0.292, -0.217] | 0.252 | VOID |
| 50 | -0.746 [-0.795, -0.694] | -0.587 [-0.626, -0.547] | -0.224 [-0.274, -0.172] | 0.147 | VOID |
The preregistered VOID marking, cell by cell
A cell whose deal rate falls more than 0.10 below base has its closure-conditional metrics stamped uninterpretable by construction — not because they look bad, but because a conditional average over a shrinking, self-selected set of closed deals is not a fairness measurement.
| cell | closure-conditional metrics | why |
|---|---|---|
| divide_dollar | interpretable | deal rate held within the viability floor |
| primary | VOID | deal rate -0.226 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| prose | VOID | deal rate -0.224 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| rationaltable | interpretable | deal rate held within the viability floor |
| story_abstract | VOID | deal rate -0.285 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| story_datacenter | VOID | deal rate -0.308 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| story_festival | VOID | deal rate -0.265 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| ultimatum | interpretable | deal rate held within the viability floor |
Coverage
prose: 295 of 480 episodes scored (61%) — 296 episodes are on disk today, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle itstory_abstract: 284 of 480 episodes scored (59%) — 285 episodes are on disk today, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle itstory_datacenter: 416 of 480 episodes scored (87%) — 417 episodes are on disk today, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle it
Every evaluated cell
Cells evaluated at this rung: divide_dollar, primary, prose, rationaltable, story_abstract, story_datacenter, story_festival, ultimatum.
divide_dollar
canonical family holdout: a deterministic divide-the-dollar preset, never trained on — 15 baseline vs 15 trained episodes.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| normalized Nash welfare (unconditional) | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| below-threshold rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| at-cap rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| among-IR rate | 0.333 | 0.000 | -0.600 [-1.000, -0.333] | 5 pairs / 2 clusters |
| NNW among IR deals | 1.000 | -- | -- | -- |
| Gini | 0.356 | 0.659 | +0.389 [+0.278, +0.556] | 5 pairs / 2 clusters |
| worst-off share | 0.083 | 0.000 | -0.167 [-0.292, -0.083] | 5 pairs / 2 clusters |
| max share | 0.617 | 0.989 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 30 episode transcripts for this cell →
primary
the held-out scorable bank -- 24 unseen parameter sets x {full, private} x 10 seeds, the arm's preregistered primary — 960 baseline vs 960 trained episodes.
Closure-conditional metrics in this cell are VOID. deal rate -0.226 is more than 0.1 below base; closure-conditional metrics are not interpretable The struck-through rows below are printed for the record and must not be read as fairness results. Note that max share and worst-off share are averaged over closed deals only, so in a voided cell they too describe a self-selected subset — they are the terms that detect extraction, not unbiased estimates of it.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.893 | 0.667 | -0.226 [-0.278, -0.174] | 960 pairs / 48 clusters |
| normalized Nash welfare (unconditional) | 0.686 | 0.497 | -0.189 [-0.239, -0.140] | 960 pairs / 48 clusters |
| below-threshold rate | 0.299 | 0.081 | -0.218 [-0.254, -0.181] | 960 pairs / 48 clusters |
| at-cap rate | 0.396 | 0.327 | -0.069 [-0.118, -0.018] | 960 pairs / 48 clusters |
| among-IR rate | 0.341 | 0.325 | -0.016 [-0.069, +0.036] | 960 pairs / 48 clusters |
| NNW among IR deals | 0.849 | 0.785 | -0.049 [-0.090, -0.016] | 175 pairs / 40 clusters |
| Gini | 0.330 | 0.255 | -0.076 [-0.096, -0.054] | 960 pairs / 48 clusters |
| worst-off share | -0.028 | 0.011 | +0.033 [+0.019, +0.050] | 580 pairs / 48 clusters |
| max share | 0.365 | 0.350 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 1920 episode transcripts for this cell →
prose
the north-star transfer cell: the same games with no machine-readable score sheet, only prose — 480 baseline vs 295 trained episodes.
Closure-conditional metrics in this cell are VOID. deal rate -0.224 is more than 0.1 below base; closure-conditional metrics are not interpretable The struck-through rows below are printed for the record and must not be read as fairness results. Note that max share and worst-off share are averaged over closed deals only, so in a voided cell they too describe a self-selected subset — they are the terms that detect extraction, not unbiased estimates of it.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.735 | 0.522 | -0.224 [-0.297, -0.151] | 295 pairs / 30 clusters |
| normalized Nash welfare (unconditional) | 0.611 | 0.414 | -0.203 [-0.269, -0.136] | 295 pairs / 30 clusters |
| below-threshold rate | 0.246 | 0.010 | -0.268 [-0.335, -0.203] | 295 pairs / 30 clusters |
| at-cap rate | 0.510 | 0.495 | +0.037 [-0.034, +0.110] | 295 pairs / 30 clusters |
| among-IR rate | 0.256 | 0.207 | -0.031 [-0.107, +0.045] | 295 pairs / 30 clusters |
| NNW among IR deals | 0.871 | 0.829 | -0.010 [-0.056, +0.032] | 25 pairs / 14 clusters |
| Gini | 0.267 | 0.215 | -0.056 [-0.082, -0.029] | 295 pairs / 30 clusters |
| worst-off share | -0.013 | 0.015 | +0.033 [+0.020, +0.046] | 118 pairs / 30 clusters |
| max share | 0.339 | 0.359 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 776 episode transcripts for this cell →
rationaltable
the exploitability guard -- the trained seat against five computable rational agents — 960 baseline vs 960 trained episodes.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.720 | 0.708 | -0.011 [-0.040, +0.017] | 960 pairs / 48 clusters |
| normalized Nash welfare (unconditional) | 0.613 | 0.606 | -0.007 [-0.031, +0.017] | 960 pairs / 48 clusters |
| below-threshold rate | 0.013 | 0.004 | -0.008 [-0.019, +0.000] | 960 pairs / 48 clusters |
| at-cap rate | 0.951 | 0.897 | -0.054 [-0.079, -0.033] | 960 pairs / 48 clusters |
| among-IR rate | 0.511 | 0.527 | +0.016 [-0.014, +0.044] | 960 pairs / 48 clusters |
| NNW among IR deals | 0.869 | 0.876 | +0.001 [-0.002, +0.005] | 454 pairs / 38 clusters |
| Gini | 0.241 | 0.235 | -0.006 [-0.015, +0.003] | 960 pairs / 48 clusters |
| worst-off share | 0.032 | 0.035 | +0.002 [-0.000, +0.004] | 639 pairs / 47 clusters |
| max share | 0.316 | 0.314 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 1920 episode transcripts for this cell →
story_abstract
story transfer with abstract issue/option labels (no domain skin) — 480 baseline vs 284 trained episodes.
Closure-conditional metrics in this cell are VOID. deal rate -0.285 is more than 0.1 below base; closure-conditional metrics are not interpretable The struck-through rows below are printed for the record and must not be read as fairness results. Note that max share and worst-off share are averaged over closed deals only, so in a voided cell they too describe a self-selected subset — they are the terms that detect extraction, not unbiased estimates of it.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.802 | 0.514 | -0.285 [-0.362, -0.207] | 284 pairs / 29 clusters |
| normalized Nash welfare (unconditional) | 0.611 | 0.383 | -0.217 [-0.267, -0.167] | 284 pairs / 29 clusters |
| below-threshold rate | 0.260 | 0.032 | -0.246 [-0.313, -0.179] | 284 pairs / 29 clusters |
| at-cap rate | 0.325 | 0.507 | +0.176 [+0.106, +0.250] | 284 pairs / 29 clusters |
| among-IR rate | 0.319 | 0.261 | -0.056 [-0.110, -0.004] | 284 pairs / 29 clusters |
| NNW among IR deals | 0.837 | 0.798 | -0.045 [-0.084, -0.007] | 52 pairs / 17 clusters |
| Gini | 0.304 | 0.202 | -0.098 [-0.132, -0.064] | 284 pairs / 29 clusters |
| worst-off share | -0.029 | 0.009 | +0.018 [-0.000, +0.040] | 124 pairs / 27 clusters |
| max share | 0.368 | 0.360 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 765 episode transcripts for this cell →
story_datacenter
story transfer with the datacenter skin — 480 baseline vs 416 trained episodes.
Closure-conditional metrics in this cell are VOID. deal rate -0.308 is more than 0.1 below base; closure-conditional metrics are not interpretable The struck-through rows below are printed for the record and must not be read as fairness results. Note that max share and worst-off share are averaged over closed deals only, so in a voided cell they too describe a self-selected subset — they are the terms that detect extraction, not unbiased estimates of it.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.808 | 0.510 | -0.308 [-0.382, -0.238] | 416 pairs / 42 clusters |
| normalized Nash welfare (unconditional) | 0.616 | 0.411 | -0.210 [-0.270, -0.156] | 416 pairs / 42 clusters |
| below-threshold rate | 0.246 | 0.017 | -0.219 [-0.276, -0.161] | 416 pairs / 42 clusters |
| at-cap rate | 0.325 | 0.476 | +0.168 [+0.094, +0.243] | 416 pairs / 42 clusters |
| among-IR rate | 0.331 | 0.315 | -0.029 [-0.075, +0.014] | 416 pairs / 42 clusters |
| NNW among IR deals | 0.832 | 0.824 | -0.003 [-0.022, +0.019] | 74 pairs / 26 clusters |
| Gini | 0.295 | 0.185 | -0.113 [-0.143, -0.085] | 416 pairs / 42 clusters |
| worst-off share | -0.016 | 0.019 | +0.033 [+0.012, +0.058] | 178 pairs / 40 clusters |
| max share | 0.360 | 0.342 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 897 episode transcripts for this cell →
story_festival
story transfer with the festival skin — 480 baseline vs 480 trained episodes.
Closure-conditional metrics in this cell are VOID. deal rate -0.265 is more than 0.1 below base; closure-conditional metrics are not interpretable The struck-through rows below are printed for the record and must not be read as fairness results. Note that max share and worst-off share are averaged over closed deals only, so in a voided cell they too describe a self-selected subset — they are the terms that detect extraction, not unbiased estimates of it.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.835 | 0.571 | -0.265 [-0.321, -0.210] | 480 pairs / 48 clusters |
| normalized Nash welfare (unconditional) | 0.649 | 0.464 | -0.185 [-0.236, -0.139] | 480 pairs / 48 clusters |
| below-threshold rate | 0.208 | 0.008 | -0.200 [-0.233, -0.163] | 480 pairs / 48 clusters |
| at-cap rate | 0.304 | 0.421 | +0.117 [+0.056, +0.177] | 480 pairs / 48 clusters |
| among-IR rate | 0.342 | 0.317 | -0.025 [-0.083, +0.033] | 480 pairs / 48 clusters |
| NNW among IR deals | 0.821 | 0.850 | +0.024 [-0.007, +0.053] | 85 pairs / 31 clusters |
| Gini | 0.304 | 0.216 | -0.088 [-0.111, -0.065] | 480 pairs / 48 clusters |
| worst-off share | -0.005 | 0.021 | +0.019 [+0.010, +0.030] | 239 pairs / 47 clusters |
| max share | 0.344 | 0.340 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 960 episode transcripts for this cell →
ultimatum
canonical family holdout: a deterministic ultimatum preset, never trained on — 15 baseline vs 15 trained episodes.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| normalized Nash welfare (unconditional) | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| below-threshold rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| at-cap rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| among-IR rate | 0.533 | 0.000 | -0.600 [-1.000, -0.500] | 5 pairs / 3 clusters |
| NNW among IR deals | 1.000 | -- | -- | -- |
| Gini | 0.233 | 0.500 | +0.300 [+0.250, +0.500] | 5 pairs / 3 clusters |
| worst-off share | 0.267 | 0.000 | -0.300 [-0.500, -0.250] | 5 pairs / 3 clusters |
| max share | 0.733 | 1.000 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 30 episode transcripts for this cell →
Provenance
- Research note (the owning record, preregistration + amendments): 0028 — Fairness-GRPO v2
- Week-notes narrative for the whole arc: fairness-grpo-v2.md
- Training run on wandb: grpo_v2_lam1 and its resume leg
grpo_v2_lam1_resume25 (group
fairness-grpo-v2) - Public dataset: 2026.RA.Fairness-GRPO
- Frozen eval summary this page is computed from:
eval_ckpt25.json(trained arm keylam1_checkpoint-25, deal-rate viability floor 0.10) - Episode transcripts published under this checkpoint: 7298 across 8 cells, re-rendered from the stored episode JSONs with the current viewer.