← the two attractors · all runs
attractor 1 — boundary extractionCheckpoint 25: it closes every canonical game, and takes the whole pie
On the ultimatum family holdout this policy closes 1.000 of its episodes, pushes 0.000 of them below a walk-away, and proposes a 100/0 split. The responder is placed exactly on its threshold: it clears IR by nothing at all, which is why the among-IR rate reads 0.000 while the below-threshold rate reads a perfect 0.000 too. It is not a tendency but a fixed point: every published ultimatum episode lands on the same package, where the base model split evenly in 8 of its 15. Only the share term sees what happened. On the held-out scorable bank the same policy has already lost 0.226 of its deal rate, so every closure-conditional fairness number there is VOID.
The held-out primary bank
24 unseen parameter sets × {full, private} × 10 seeds, paired on (instance, seed, arm); intervals are 95% cluster bootstraps over parameter sets. The deal rate is the number that stays interpretable whatever else happens, so it is printed beside every conditional metric.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.893 | 0.667 | -0.226 [-0.278, -0.174] | 960 pairs / 48 clusters |
| normalized Nash welfare (unconditional) | 0.686 | 0.497 | -0.189 [-0.239, -0.140] | 960 pairs / 48 clusters |
| below-threshold rate | 0.299 | 0.081 | -0.218 [-0.254, -0.181] | 960 pairs / 48 clusters |
| at-cap rate | 0.396 | 0.327 | -0.069 [-0.118, -0.018] | 960 pairs / 48 clusters |
| among-IR rate | 0.341 | 0.325 | -0.016 [-0.069, +0.036] | 960 pairs / 48 clusters |
| NNW among IR deals | 0.849 | 0.785 | -0.049 [-0.090, -0.016] | 175 pairs / 40 clusters |
| Gini | 0.330 | 0.255 | -0.076 [-0.096, -0.054] | 960 pairs / 48 clusters |
| worst-off share | -0.028 | 0.011 | +0.033 [+0.019, +0.050] | 580 pairs / 48 clusters |
| max share | 0.365 | 0.350 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Where the run sits on the ladder
All seven evaluated rungs of the λ=1.0 arm, held-out primary bank, trained minus base. Held-out deal rate falls monotonically and never recovers, while the below-threshold rate stays favourable throughout — the deal-suppression signature, stated as plainly as it can be.
| checkpoint | Δ deal rate | Δ NNW (unconditional) | Δ below-threshold | trained deal rate | closure-conditional |
|---|---|---|---|---|---|
| 5 | -0.052 [-0.081, -0.021] | -0.044 [-0.072, -0.013] | -0.104 [-0.138, -0.071] | 0.841 | interpretable |
| 10 | -0.143 [-0.184, -0.099] | -0.132 [-0.170, -0.092] | -0.207 [-0.243, -0.172] | 0.750 | VOID |
| 15 | -0.147 [-0.197, -0.095] | -0.156 [-0.202, -0.108] | -0.185 [-0.225, -0.144] | 0.746 | VOID |
| 25 | -0.226 [-0.278, -0.174] | -0.189 [-0.239, -0.140] | -0.218 [-0.254, -0.181] | 0.667 | VOID |
| 40 | -0.453 [-0.518, -0.385] | -0.348 [-0.404, -0.290] | -0.279 [-0.317, -0.242] | 0.440 | VOID |
| 45 | -0.641 [-0.682, -0.597] | -0.493 [-0.530, -0.453] | -0.254 [-0.292, -0.217] | 0.252 | VOID |
| 50 | -0.746 [-0.795, -0.694] | -0.587 [-0.626, -0.547] | -0.224 [-0.274, -0.172] | 0.147 | VOID |
The preregistered VOID marking, cell by cell
A cell whose deal rate falls more than 0.10 below base has its closure-conditional metrics stamped uninterpretable by construction — not because they look bad, but because a conditional average over a shrinking, self-selected set of closed deals is not a fairness measurement.
| cell | closure-conditional metrics | why |
|---|---|---|
| divide_dollar | interpretable | deal rate held within the viability floor |
| primary | VOID | deal rate -0.226 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| prose | VOID | deal rate -0.231 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| rationaltable | interpretable | deal rate held within the viability floor |
| story_abstract | VOID | deal rate -0.258 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| story_datacenter | VOID | deal rate -0.320 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| story_festival | VOID | deal rate -0.244 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| ultimatum | interpretable | deal rate held within the viability floor |
Coverage
prose: 247 of 480 episodes scored (51%) — 296 episodes are on disk today, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle itrationaltable: 820 of 960 episodes scored (85%) — 960 episodes are on disk today, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle itstory_abstract: 233 of 480 episodes scored (49%) — 285 episodes are on disk today, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle itstory_datacenter: 200 of 480 episodes scored (42%) — 417 episodes are on disk today, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle itstory_festival: 168 of 480 episodes scored (35%) — 480 episodes are on disk today, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle it
Every evaluated cell
Cells evaluated at this rung: divide_dollar, primary, prose, rationaltable, story_abstract, story_datacenter, story_festival, ultimatum.
divide_dollar
canonical family holdout: a deterministic divide-the-dollar preset, never trained on — 15 baseline vs 15 trained episodes.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| normalized Nash welfare (unconditional) | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| below-threshold rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| at-cap rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| among-IR rate | 0.333 | 0.000 | -0.600 [-1.000, -0.333] | 5 pairs / 2 clusters |
| NNW among IR deals | 1.000 | -- | -- | -- |
| Gini | 0.356 | 0.659 | +0.389 [+0.278, +0.556] | 5 pairs / 2 clusters |
| worst-off share | 0.083 | 0.000 | -0.167 [-0.292, -0.083] | 5 pairs / 2 clusters |
| max share | 0.617 | 0.989 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 30 episode transcripts for this cell →
primary
the held-out scorable bank -- 24 unseen parameter sets x {full, private} x 10 seeds, the arm's preregistered primary — 960 baseline vs 960 trained episodes.
Closure-conditional metrics in this cell are VOID. deal rate -0.226 is more than 0.1 below base; closure-conditional metrics are not interpretable The struck-through rows below are printed for the record and must not be read as fairness results. Note that max share and worst-off share are averaged over closed deals only, so in a voided cell they too describe a self-selected subset — they are the terms that detect extraction, not unbiased estimates of it.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.893 | 0.667 | -0.226 [-0.278, -0.174] | 960 pairs / 48 clusters |
| normalized Nash welfare (unconditional) | 0.686 | 0.497 | -0.189 [-0.239, -0.140] | 960 pairs / 48 clusters |
| below-threshold rate | 0.299 | 0.081 | -0.218 [-0.254, -0.181] | 960 pairs / 48 clusters |
| at-cap rate | 0.396 | 0.327 | -0.069 [-0.118, -0.018] | 960 pairs / 48 clusters |
| among-IR rate | 0.341 | 0.325 | -0.016 [-0.069, +0.036] | 960 pairs / 48 clusters |
| NNW among IR deals | 0.849 | 0.785 | -0.049 [-0.090, -0.016] | 175 pairs / 40 clusters |
| Gini | 0.330 | 0.255 | -0.076 [-0.096, -0.054] | 960 pairs / 48 clusters |
| worst-off share | -0.028 | 0.011 | +0.033 [+0.019, +0.050] | 580 pairs / 48 clusters |
| max share | 0.365 | 0.350 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 1920 episode transcripts for this cell →
prose
the north-star transfer cell: the same games with no machine-readable score sheet, only prose — 480 baseline vs 247 trained episodes.
Closure-conditional metrics in this cell are VOID. deal rate -0.231 is more than 0.1 below base; closure-conditional metrics are not interpretable The struck-through rows below are printed for the record and must not be read as fairness results. Note that max share and worst-off share are averaged over closed deals only, so in a voided cell they too describe a self-selected subset — they are the terms that detect extraction, not unbiased estimates of it.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.735 | 0.547 | -0.231 [-0.316, -0.140] | 247 pairs / 25 clusters |
| normalized Nash welfare (unconditional) | 0.611 | 0.430 | -0.217 [-0.296, -0.135] | 247 pairs / 25 clusters |
| below-threshold rate | 0.246 | 0.008 | -0.279 [-0.360, -0.204] | 247 pairs / 25 clusters |
| at-cap rate | 0.510 | 0.474 | +0.045 [-0.036, +0.121] | 247 pairs / 25 clusters |
| among-IR rate | 0.256 | 0.198 | -0.028 [-0.113, +0.053] | 247 pairs / 25 clusters |
| NNW among IR deals | 0.871 | 0.832 | -0.003 [-0.047, +0.035] | 19 pairs / 10 clusters |
| Gini | 0.267 | 0.229 | -0.056 [-0.089, -0.022] | 247 pairs / 25 clusters |
| worst-off share | -0.013 | 0.014 | +0.032 [+0.019, +0.044] | 106 pairs / 24 clusters |
| max share | 0.339 | 0.363 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 776 episode transcripts for this cell →
rationaltable
the exploitability guard -- the trained seat against five computable rational agents — 960 baseline vs 820 trained episodes.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.720 | 0.716 | -0.023 [-0.056, +0.007] | 820 pairs / 41 clusters |
| normalized Nash welfare (unconditional) | 0.613 | 0.609 | -0.017 [-0.046, +0.008] | 820 pairs / 41 clusters |
| below-threshold rate | 0.013 | 0.005 | -0.010 [-0.022, +0.000] | 820 pairs / 41 clusters |
| at-cap rate | 0.951 | 0.893 | -0.052 [-0.079, -0.028] | 820 pairs / 41 clusters |
| among-IR rate | 0.511 | 0.560 | +0.006 [-0.028, +0.038] | 820 pairs / 41 clusters |
| NNW among IR deals | 0.869 | 0.874 | +0.002 [-0.002, +0.005] | 419 pairs / 34 clusters |
| Gini | 0.241 | 0.239 | -0.008 [-0.018, +0.003] | 820 pairs / 41 clusters |
| worst-off share | 0.032 | 0.037 | +0.001 [-0.001, +0.004] | 557 pairs / 41 clusters |
| max share | 0.316 | 0.316 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 1920 episode transcripts for this cell →
story_abstract
story transfer with abstract issue/option labels (no domain skin) — 480 baseline vs 233 trained episodes.
Closure-conditional metrics in this cell are VOID. deal rate -0.258 is more than 0.1 below base; closure-conditional metrics are not interpretable The struck-through rows below are printed for the record and must not be read as fairness results. Note that max share and worst-off share are averaged over closed deals only, so in a voided cell they too describe a self-selected subset — they are the terms that detect extraction, not unbiased estimates of it.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.802 | 0.545 | -0.258 [-0.342, -0.170] | 233 pairs / 24 clusters |
| normalized Nash welfare (unconditional) | 0.611 | 0.405 | -0.197 [-0.252, -0.138] | 233 pairs / 24 clusters |
| below-threshold rate | 0.260 | 0.030 | -0.219 [-0.288, -0.155] | 233 pairs / 24 clusters |
| at-cap rate | 0.325 | 0.481 | +0.150 [+0.079, +0.223] | 233 pairs / 24 clusters |
| among-IR rate | 0.319 | 0.275 | -0.069 [-0.133, -0.008] | 233 pairs / 24 clusters |
| NNW among IR deals | 0.837 | 0.806 | -0.045 [-0.086, -0.001] | 47 pairs / 16 clusters |
| Gini | 0.304 | 0.214 | -0.087 [-0.124, -0.049] | 233 pairs / 24 clusters |
| worst-off share | -0.029 | 0.009 | +0.011 [-0.006, +0.031] | 107 pairs / 22 clusters |
| max share | 0.368 | 0.358 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 765 episode transcripts for this cell →
story_datacenter
story transfer with the datacenter skin — 480 baseline vs 200 trained episodes.
Closure-conditional metrics in this cell are VOID. deal rate -0.320 is more than 0.1 below base; closure-conditional metrics are not interpretable The struck-through rows below are printed for the record and must not be read as fairness results. Note that max share and worst-off share are averaged over closed deals only, so in a voided cell they too describe a self-selected subset — they are the terms that detect extraction, not unbiased estimates of it.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.808 | 0.495 | -0.320 [-0.420, -0.230] | 200 pairs / 20 clusters |
| normalized Nash welfare (unconditional) | 0.616 | 0.404 | -0.213 [-0.290, -0.143] | 200 pairs / 20 clusters |
| below-threshold rate | 0.246 | 0.010 | -0.210 [-0.290, -0.140] | 200 pairs / 20 clusters |
| at-cap rate | 0.325 | 0.510 | +0.220 [+0.140, +0.310] | 200 pairs / 20 clusters |
| among-IR rate | 0.331 | 0.360 | -0.050 [-0.115, +0.010] | 200 pairs / 20 clusters |
| NNW among IR deals | 0.832 | 0.831 | -0.011 [-0.039, +0.021] | 42 pairs / 15 clusters |
| Gini | 0.295 | 0.176 | -0.120 [-0.160, -0.082] | 200 pairs / 20 clusters |
| worst-off share | -0.016 | 0.031 | +0.042 [+0.015, +0.085] | 79 pairs / 20 clusters |
| max share | 0.360 | 0.330 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 897 episode transcripts for this cell →
story_festival
story transfer with the festival skin — 480 baseline vs 168 trained episodes.
Closure-conditional metrics in this cell are VOID. deal rate -0.244 is more than 0.1 below base; closure-conditional metrics are not interpretable The struck-through rows below are printed for the record and must not be read as fairness results. Note that max share and worst-off share are averaged over closed deals only, so in a voided cell they too describe a self-selected subset — they are the terms that detect extraction, not unbiased estimates of it.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.835 | 0.571 | -0.244 [-0.327, -0.163] | 168 pairs / 17 clusters |
| normalized Nash welfare (unconditional) | 0.649 | 0.461 | -0.181 [-0.266, -0.109] | 168 pairs / 17 clusters |
| below-threshold rate | 0.208 | 0.006 | -0.155 [-0.217, -0.095] | 168 pairs / 17 clusters |
| at-cap rate | 0.304 | 0.423 | +0.125 [+0.030, +0.223] | 168 pairs / 17 clusters |
| among-IR rate | 0.342 | 0.292 | -0.101 [-0.190, -0.012] | 168 pairs / 17 clusters |
| NNW among IR deals | 0.821 | 0.878 | +0.041 [-0.022, +0.092] | 30 pairs / 10 clusters |
| Gini | 0.304 | 0.232 | -0.068 [-0.108, -0.027] | 168 pairs / 17 clusters |
| worst-off share | -0.005 | 0.015 | +0.007 [-0.003, +0.018] | 83 pairs / 16 clusters |
| max share | 0.344 | 0.351 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 960 episode transcripts for this cell →
ultimatum
canonical family holdout: a deterministic ultimatum preset, never trained on — 15 baseline vs 15 trained episodes.
| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| normalized Nash welfare (unconditional) | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| below-threshold rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| at-cap rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| among-IR rate | 0.533 | 0.000 | -0.600 [-1.000, -0.500] | 5 pairs / 3 clusters |
| NNW among IR deals | 1.000 | -- | -- | -- |
| Gini | 0.233 | 0.500 | +0.300 [+0.250, +0.500] | 5 pairs / 3 clusters |
| worst-off share | 0.267 | 0.000 | -0.300 [-0.500, -0.250] | 5 pairs / 3 clusters |
| max share | 0.733 | 1.000 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |
Browse all 30 episode transcripts for this cell →
Provenance
- Research note (the owning record, preregistration + amendments): 0028 — Fairness-GRPO v2
- Week-notes narrative for the whole arc: fairness-grpo-v2.md
- Training run on wandb: grpo_v2_lam1 and its resume leg
grpo_v2_lam1_resume25 (group
fairness-grpo-v2) - Public dataset: 2026.RA.Fairness-GRPO
- Frozen eval summary this page is computed from:
eval_ckpt25.json(trained arm keylam1_checkpoint-25, deal-rate viability floor 0.10) - Episode transcripts published under this checkpoint: 7298 across 8 cells, re-rendered from the stored episode JSONs with the current viewer.