<!-- [fix: rational_agents orig (results review)] 2026-08-10 — generated by experiments/rational_agents/build_grpo_v2_checkpoint_hub.py from eval_ckpt25.json; do not hand-edit. -->

# Fairness-GRPO v2, λ=1.0 — checkpoint 25

**ATTRACTOR 1 — BOUNDARY EXTRACTION.** Checkpoint 25: it closes every canonical game, and takes the whole pie

On the ultimatum family holdout this policy closes **1.000** of its episodes, pushes **0.000** of them below a walk-away, and proposes a **100/0** split. The responder is placed *exactly on* its threshold: it clears IR by nothing at all, which is why the among-IR rate reads 0.000 while the below-threshold rate reads a perfect 0.000 too. It is not a tendency but a fixed point: every published ultimatum episode lands on the same package, where the base model split evenly in 8 of its 15. Only the share term sees what happened. On the held-out scorable bank the same policy has already lost 0.226 of its deal rate, so every closure-conditional fairness number there is VOID.

## Held-out primary bank

| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.893 | 0.667 | -0.226 [-0.278, -0.174] | 960 pairs / 48 clusters |
| normalized Nash welfare (unconditional) | 0.686 | 0.497 | -0.189 [-0.239, -0.140] | 960 pairs / 48 clusters |
| below-threshold rate | 0.299 | 0.081 | -0.218 [-0.254, -0.181] | 960 pairs / 48 clusters |
| at-cap rate | 0.396 | 0.327 | -0.069 [-0.118, -0.018] | 960 pairs / 48 clusters |
| among-IR rate | 0.341 | 0.325 | -0.016 [-0.069, +0.036] | 960 pairs / 48 clusters |
| NNW among IR deals | ~~0.849~~ | ~~0.785~~ | ~~-0.049 [-0.090, -0.016]~~ | ~~175 pairs / 40 clusters~~ |
| Gini | ~~0.330~~ | ~~0.255~~ | ~~-0.076 [-0.096, -0.054]~~ | ~~960 pairs / 48 clusters~~ |
| worst-off share | ~~-0.028~~ | ~~0.011~~ | ~~+0.033 [+0.019, +0.050]~~ | ~~580 pairs / 48 clusters~~ |
| max share | 0.365 | 0.350 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |

Struck-through numbers are the preregistered VOID marking: the cell's deal rate fell more than 0.10 below base, so its closure-conditional metrics are uninterpretable by construction. `max share` and `worst-off share` are likewise averaged over closed deals only: in a voided cell they describe a self-selected subset, and are reported because they are the terms that *detect* extraction, not as unbiased estimates of it.

## The λ=1.0 ladder (held-out primary, trained − base)

| checkpoint | Δ deal rate | Δ NNW (unconditional) | Δ below-threshold | trained deal rate | closure-conditional |
|---|---|---|---|---|---|
| 5 | -0.052 [-0.081, -0.021] | -0.044 [-0.072, -0.013] | -0.104 [-0.138, -0.071] | 0.841 | interpretable |
| 10 | -0.143 [-0.184, -0.099] | -0.132 [-0.170, -0.092] | -0.207 [-0.243, -0.172] | 0.750 | VOID |
| 15 | -0.147 [-0.197, -0.095] | -0.156 [-0.202, -0.108] | -0.185 [-0.225, -0.144] | 0.746 | VOID |
| 25 | -0.226 [-0.278, -0.174] | -0.189 [-0.239, -0.140] | -0.218 [-0.254, -0.181] | 0.667 | VOID |
| 40 | -0.453 [-0.518, -0.385] | -0.348 [-0.404, -0.290] | -0.279 [-0.317, -0.242] | 0.440 | VOID |
| 45 | -0.641 [-0.682, -0.597] | -0.493 [-0.530, -0.453] | -0.254 [-0.292, -0.217] | 0.252 | VOID |
| 50 | -0.746 [-0.795, -0.694] | -0.587 [-0.626, -0.547] | -0.224 [-0.274, -0.172] | 0.147 | VOID |

## Per-cell VOID status

| cell | closure-conditional metrics | why |
|---|---|---|
| divide_dollar | interpretable | deal rate held within the viability floor |
| primary | VOID | deal rate -0.226 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| prose | VOID | deal rate -0.231 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| rationaltable | interpretable | deal rate held within the viability floor |
| story_abstract | VOID | deal rate -0.258 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| story_datacenter | VOID | deal rate -0.320 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| story_festival | VOID | deal rate -0.244 is more than 0.1 below base; closure-conditional metrics are not interpretable |
| ultimatum | interpretable | deal rate held within the viability floor |

## Partial cells

These cells scored the trained arm on fewer than 95% of the baseline's episodes. Completion bias is unbounded — read them as mid-flight samples, not results.

- `prose`: 247 of 480 episodes scored (51%) — **296 episodes are on disk today**, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle it
- `rationaltable`: 820 of 960 episodes scored (85%) — **960 episodes are on disk today**, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle it
- `story_abstract`: 233 of 480 episodes scored (49%) — **285 episodes are on disk today**, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle it
- `story_datacenter`: 200 of 480 episodes scored (42%) — **417 episodes are on disk today**, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle it
- `story_festival`: 168 of 480 episodes scored (35%) — **480 episodes are on disk today**, so this row is a snapshot of a bank that has since grown; re-analysis, not re-running, is what would settle it

## Every evaluated cell

### `divide_dollar`

canonical family holdout: a deterministic divide-the-dollar preset, never trained on
(15 baseline vs 15 trained episodes; closure-conditional metrics **interpretable**)

| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| normalized Nash welfare (unconditional) | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| below-threshold rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| at-cap rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 2 clusters |
| among-IR rate | 0.333 | 0.000 | -0.600 [-1.000, -0.333] | 5 pairs / 2 clusters |
| NNW among IR deals | 1.000 | -- | -- | -- |
| Gini | 0.356 | 0.659 | +0.389 [+0.278, +0.556] | 5 pairs / 2 clusters |
| worst-off share | 0.083 | 0.000 | -0.167 [-0.292, -0.083] | 5 pairs / 2 clusters |
| max share | 0.617 | 0.989 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |

[Browse all 30 episode transcripts for this cell](episodes/divide_dollar/index.html). Source run directories: `baseline` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_divide_dollar`; `trained` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_divide_dollar`.

### `primary`

the held-out scorable bank -- 24 unseen parameter sets x {full, private} x 10 seeds, the arm's preregistered primary
(960 baseline vs 960 trained episodes; closure-conditional metrics **VOID**)

| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.893 | 0.667 | -0.226 [-0.278, -0.174] | 960 pairs / 48 clusters |
| normalized Nash welfare (unconditional) | 0.686 | 0.497 | -0.189 [-0.239, -0.140] | 960 pairs / 48 clusters |
| below-threshold rate | 0.299 | 0.081 | -0.218 [-0.254, -0.181] | 960 pairs / 48 clusters |
| at-cap rate | 0.396 | 0.327 | -0.069 [-0.118, -0.018] | 960 pairs / 48 clusters |
| among-IR rate | 0.341 | 0.325 | -0.016 [-0.069, +0.036] | 960 pairs / 48 clusters |
| NNW among IR deals | ~~0.849~~ | ~~0.785~~ | ~~-0.049 [-0.090, -0.016]~~ | ~~175 pairs / 40 clusters~~ |
| Gini | ~~0.330~~ | ~~0.255~~ | ~~-0.076 [-0.096, -0.054]~~ | ~~960 pairs / 48 clusters~~ |
| worst-off share | ~~-0.028~~ | ~~0.011~~ | ~~+0.033 [+0.019, +0.050]~~ | ~~580 pairs / 48 clusters~~ |
| max share | 0.365 | 0.350 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |

[Browse all 1920 episode transcripts for this cell](episodes/primary/index.html). Source run directories: `baseline` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_primary_s0, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_primary_s1, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_primary_s2, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_primary_s3, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_primary_s4, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_primary_s5, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_primary_s6, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_primary_s7, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_primary_s8, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_primary_s9`; `trained` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_primary_s0, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_primary_s1, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_primary_s2, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_primary_s3, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_primary_s4, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_primary_s5, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_primary_s6, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_primary_s7, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_primary_s8, /nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_primary_s9`.

### `prose`

the north-star transfer cell: the same games with no machine-readable score sheet, only prose
(480 baseline vs 247 trained episodes; closure-conditional metrics **VOID**)

| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.735 | 0.547 | -0.231 [-0.316, -0.140] | 247 pairs / 25 clusters |
| normalized Nash welfare (unconditional) | 0.611 | 0.430 | -0.217 [-0.296, -0.135] | 247 pairs / 25 clusters |
| below-threshold rate | 0.246 | 0.008 | -0.279 [-0.360, -0.204] | 247 pairs / 25 clusters |
| at-cap rate | 0.510 | 0.474 | +0.045 [-0.036, +0.121] | 247 pairs / 25 clusters |
| among-IR rate | 0.256 | 0.198 | -0.028 [-0.113, +0.053] | 247 pairs / 25 clusters |
| NNW among IR deals | ~~0.871~~ | ~~0.832~~ | ~~-0.003 [-0.047, +0.035]~~ | ~~19 pairs / 10 clusters~~ |
| Gini | ~~0.267~~ | ~~0.229~~ | ~~-0.056 [-0.089, -0.022]~~ | ~~247 pairs / 25 clusters~~ |
| worst-off share | ~~-0.013~~ | ~~0.014~~ | ~~+0.032 [+0.019, +0.044]~~ | ~~106 pairs / 24 clusters~~ |
| max share | 0.339 | 0.363 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |

[Browse all 776 episode transcripts for this cell](episodes/prose/index.html). Source run directories: `baseline` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_prose`; `trained` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_prose`.

### `rationaltable`

the exploitability guard -- the trained seat against five computable rational agents
(960 baseline vs 820 trained episodes; closure-conditional metrics **interpretable**)

| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.720 | 0.716 | -0.023 [-0.056, +0.007] | 820 pairs / 41 clusters |
| normalized Nash welfare (unconditional) | 0.613 | 0.609 | -0.017 [-0.046, +0.008] | 820 pairs / 41 clusters |
| below-threshold rate | 0.013 | 0.005 | -0.010 [-0.022, +0.000] | 820 pairs / 41 clusters |
| at-cap rate | 0.951 | 0.893 | -0.052 [-0.079, -0.028] | 820 pairs / 41 clusters |
| among-IR rate | 0.511 | 0.560 | +0.006 [-0.028, +0.038] | 820 pairs / 41 clusters |
| NNW among IR deals | 0.869 | 0.874 | +0.002 [-0.002, +0.005] | 419 pairs / 34 clusters |
| Gini | 0.241 | 0.239 | -0.008 [-0.018, +0.003] | 820 pairs / 41 clusters |
| worst-off share | 0.032 | 0.037 | +0.001 [-0.001, +0.004] | 557 pairs / 41 clusters |
| max share | 0.316 | 0.316 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |

[Browse all 1920 episode transcripts for this cell](episodes/rationaltable/index.html). Source run directories: `baseline` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_rationaltable`; `trained` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_rationaltable`.

### `story_abstract`

story transfer with abstract issue/option labels (no domain skin)
(480 baseline vs 233 trained episodes; closure-conditional metrics **VOID**)

| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.802 | 0.545 | -0.258 [-0.342, -0.170] | 233 pairs / 24 clusters |
| normalized Nash welfare (unconditional) | 0.611 | 0.405 | -0.197 [-0.252, -0.138] | 233 pairs / 24 clusters |
| below-threshold rate | 0.260 | 0.030 | -0.219 [-0.288, -0.155] | 233 pairs / 24 clusters |
| at-cap rate | 0.325 | 0.481 | +0.150 [+0.079, +0.223] | 233 pairs / 24 clusters |
| among-IR rate | 0.319 | 0.275 | -0.069 [-0.133, -0.008] | 233 pairs / 24 clusters |
| NNW among IR deals | ~~0.837~~ | ~~0.806~~ | ~~-0.045 [-0.086, -0.001]~~ | ~~47 pairs / 16 clusters~~ |
| Gini | ~~0.304~~ | ~~0.214~~ | ~~-0.087 [-0.124, -0.049]~~ | ~~233 pairs / 24 clusters~~ |
| worst-off share | ~~-0.029~~ | ~~0.009~~ | ~~+0.011 [-0.006, +0.031]~~ | ~~107 pairs / 22 clusters~~ |
| max share | 0.368 | 0.358 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |

[Browse all 765 episode transcripts for this cell](episodes/story_abstract/index.html). Source run directories: `baseline` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_story_abstract`; `trained` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_story_abstract`.

### `story_datacenter`

story transfer with the datacenter skin
(480 baseline vs 200 trained episodes; closure-conditional metrics **VOID**)

| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.808 | 0.495 | -0.320 [-0.420, -0.230] | 200 pairs / 20 clusters |
| normalized Nash welfare (unconditional) | 0.616 | 0.404 | -0.213 [-0.290, -0.143] | 200 pairs / 20 clusters |
| below-threshold rate | 0.246 | 0.010 | -0.210 [-0.290, -0.140] | 200 pairs / 20 clusters |
| at-cap rate | 0.325 | 0.510 | +0.220 [+0.140, +0.310] | 200 pairs / 20 clusters |
| among-IR rate | 0.331 | 0.360 | -0.050 [-0.115, +0.010] | 200 pairs / 20 clusters |
| NNW among IR deals | ~~0.832~~ | ~~0.831~~ | ~~-0.011 [-0.039, +0.021]~~ | ~~42 pairs / 15 clusters~~ |
| Gini | ~~0.295~~ | ~~0.176~~ | ~~-0.120 [-0.160, -0.082]~~ | ~~200 pairs / 20 clusters~~ |
| worst-off share | ~~-0.016~~ | ~~0.031~~ | ~~+0.042 [+0.015, +0.085]~~ | ~~79 pairs / 20 clusters~~ |
| max share | 0.360 | 0.330 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |

[Browse all 897 episode transcripts for this cell](episodes/story_datacenter/index.html). Source run directories: `baseline` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_story_datacenter`; `trained` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_story_datacenter`.

### `story_festival`

story transfer with the festival skin
(480 baseline vs 168 trained episodes; closure-conditional metrics **VOID**)

| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 0.835 | 0.571 | -0.244 [-0.327, -0.163] | 168 pairs / 17 clusters |
| normalized Nash welfare (unconditional) | 0.649 | 0.461 | -0.181 [-0.266, -0.109] | 168 pairs / 17 clusters |
| below-threshold rate | 0.208 | 0.006 | -0.155 [-0.217, -0.095] | 168 pairs / 17 clusters |
| at-cap rate | 0.304 | 0.423 | +0.125 [+0.030, +0.223] | 168 pairs / 17 clusters |
| among-IR rate | 0.342 | 0.292 | -0.101 [-0.190, -0.012] | 168 pairs / 17 clusters |
| NNW among IR deals | ~~0.821~~ | ~~0.878~~ | ~~+0.041 [-0.022, +0.092]~~ | ~~30 pairs / 10 clusters~~ |
| Gini | ~~0.304~~ | ~~0.232~~ | ~~-0.068 [-0.108, -0.027]~~ | ~~168 pairs / 17 clusters~~ |
| worst-off share | ~~-0.005~~ | ~~0.015~~ | ~~+0.007 [-0.003, +0.018]~~ | ~~83 pairs / 16 clusters~~ |
| max share | 0.344 | 0.351 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |

[Browse all 960 episode transcripts for this cell](episodes/story_festival/index.html). Source run directories: `baseline` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_story_festival`; `trained` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_story_festival`.

### `ultimatum`

canonical family holdout: a deterministic ultimatum preset, never trained on
(15 baseline vs 15 trained episodes; closure-conditional metrics **interpretable**)

| metric | base | checkpoint 25 | trained − base [95% CI] | support |
|---|---|---|---|---|
| deal rate | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| normalized Nash welfare (unconditional) | 1.000 | 1.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| below-threshold rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| at-cap rate | 0.000 | 0.000 | +0.000 [+0.000, +0.000] | 5 pairs / 3 clusters |
| among-IR rate | 0.533 | 0.000 | -0.600 [-1.000, -0.500] | 5 pairs / 3 clusters |
| NNW among IR deals | 1.000 | -- | -- | -- |
| Gini | 0.233 | 0.500 | +0.300 [+0.250, +0.500] | 5 pairs / 3 clusters |
| worst-off share | 0.267 | 0.000 | -0.300 [-0.500, -0.250] | 5 pairs / 3 clusters |
| max share | 0.733 | 1.000 | -- | -- |
| generation failures | 0.000 | 0.000 | -- | -- |

[Browse all 30 episode transcripts for this cell](episodes/ultimatum/index.html). Source run directories: `baseline` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_baseline_ultimatum`; `trained` → `/nlp/scr/siddharth/ii_mats/rational_agents/grpov2eval_lam1_checkpoint-25_ultimatum`.

## Provenance

- Research note (owning record): [0028 — Fairness-GRPO v2](https://github.com/Sid-MB/ii_mats/blob/main/experiments/rational_agents/research-notes/0028-fairness-grpo-v2.md)
- Week-notes narrative: [fairness-grpo-v2.md](https://github.com/Sid-MB/ii_mats/blob/main/experiments/rational_agents/docs/08-07-26%20week%20notes/fairness-grpo-v2.md)
- wandb: [grpo_v2_lam1](https://wandb.ai/siddharth-stanford/rational_agents_fairness_grpo/runs/npz11gav) and its resume leg [grpo_v2_lam1_resume25](https://wandb.ai/siddharth-stanford/rational_agents_fairness_grpo)
- Public dataset: [2026.RA.Fairness-GRPO](https://huggingface.co/datasets/siddharthmb/2026.RA.Fairness-GRPO)
- Frozen eval summary: `results/fairness_grpo_v2/eval_ckpt25.json` (trained arm key `lam1_checkpoint-25`, deal-rate viability floor 0.10)
