# Five-arm fairness basket and correct-walk credit

All intervals are 95% cluster bootstraps (10000 resamples) over the 24 parameter sets, each set's five seeds resampled together. Paired contrasts join on `(instance_id, episode_seed)`.


## Fairness basket — conditional on a closed deal

Every column below is computed over closed deals only, so `deal rate` is carried beside them: an arm that closes 28 of 120 deals is describing a self-selected set of games, not the same games. Normalized coordinate z_i = (u_i − tau_i)/c_i with c_i the party's best surplus on the individually rational set (affine-invariant per party).

| arm | n | deal rate | below-thr. accept | worst-off z | worst-off share | dist-NBS | dist-KS | norm. Gini | norm. Nash welf. | max share |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| `all_llm` | 120 | 0.958 [0.925, 0.992] | 0.000 [0.000, 0.000] | 0.257 [0.206, 0.306] | 0.077 [0.064, 0.090] | 0.464 [0.357, 0.586] | 0.398 [0.316, 0.474] | 0.223 [0.200, 0.247] | 0.535 [0.458, 0.601] | 0.299 [0.287, 0.309] |
| `all_rational` | 120 | 0.233 [0.133, 0.342] | 0.000 [0.000, 0.000] | 0.147 [0.103, 0.196] | 0.052 [0.037, 0.068] | 0.600 [0.498, 0.681] | 0.590 [0.520, 0.649] | 0.261 [0.238, 0.283] | 0.408 [0.318, 0.486] | 0.314 [0.299, 0.327] |
| `one_rational` | 120 | 0.767 [0.683, 0.842] | 0.000 [0.000, 0.000] | 0.239 [0.195, 0.282] | 0.073 [0.061, 0.083] | 0.444 [0.341, 0.564] | 0.414 [0.332, 0.493] | 0.231 [0.209, 0.255] | 0.506 [0.439, 0.565] | 0.301 [0.289, 0.314] |
| `rational_advised_llm` | 120 | 0.900 [0.808, 0.975] | 0.000 [0.000, 0.000] | 0.250 [0.206, 0.294] | 0.076 [0.064, 0.087] | 0.478 [0.385, 0.581] | 0.458 [0.372, 0.535] | 0.232 [0.209, 0.253] | 0.526 [0.465, 0.581] | 0.302 [0.291, 0.313] |
| `interpreter_advised_llm` | 120 | 0.958 [0.917, 0.992] | 0.000 [0.000, 0.000] | 0.269 [0.224, 0.310] | 0.082 [0.069, 0.093] | 0.412 [0.311, 0.536] | 0.366 [0.295, 0.430] | 0.214 [0.196, 0.233] | 0.543 [0.477, 0.598] | 0.296 [0.285, 0.307] |
| `all_interpreter_advised_llm` | 120 | 0.983 [0.950, 1.000] | 0.000 [0.000, 0.000] | 0.239 [0.193, 0.286] | 0.073 [0.060, 0.085] | 0.500 [0.392, 0.633] | 0.442 [0.353, 0.524] | 0.229 [0.207, 0.250] | 0.511 [0.436, 0.576] | 0.300 [0.288, 0.311] |
| `one_oracle` | 120 | 0.917 [0.850, 0.967] | 0.000 [0.000, 0.000] | 0.201 [0.151, 0.252] | 0.059 [0.047, 0.070] | 0.550 [0.436, 0.676] | 0.524 [0.426, 0.614] | 0.266 [0.237, 0.294] | 0.472 [0.393, 0.548] | 0.314 [0.297, 0.332] |
| `all_oracle` | 120 | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | 0.166 [0.118, 0.220] | 0.048 [0.037, 0.061] | 0.631 [0.523, 0.758] | 0.549 [0.450, 0.648] | 0.285 [0.255, 0.314] | 0.406 [0.317, 0.495] | 0.319 [0.303, 0.336] |

Lower is better for below-threshold accept, dist-NBS, dist-KS, Gini, and max share; higher is better for worst-off z, worst-off share, and normalized Nash welfare. An equal split puts worst-off share at 0.200 and max share at 0.200.


### Among-deals surplus (per-party z and episode score)

Both blocks condition on a closed deal, so deal rate is carried beside them: an arm's closed set is self-selected, and a low-deal-rate arm's columns describe a different, easier subset of games. Per-party z pools party-observations (five per closed deal, no per-deal averaging; mean with an instance-clustered bootstrap CI, quartiles empirical/descriptive); the episode-score block is descriptive (the unconditional headline already carries the arm's interval).

| arm | deals | deal rate | per-party z mean [95% CI] | z median | z IQR | score mean | score median | score IQR |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `all_llm` | 115/120 | 0.958 | 0.653 [0.619, 0.691] | 0.667 | [0.433, 1.000] | 0.911 | 0.946 | [0.838, 1.000] |
| `all_rational` | 28/120 | 0.233 | 0.554 [0.512, 0.611] | 0.567 | [0.329, 0.762] | 0.809 | 0.830 | [0.699, 0.922] |
| `one_rational` | 92/120 | 0.767 | 0.639 [0.606, 0.675] | 0.657 | [0.431, 1.000] | 0.895 | 0.931 | [0.825, 1.000] |
| `rational_advised_llm` | 108/120 | 0.900 | 0.646 [0.616, 0.678] | 0.667 | [0.414, 1.000] | 0.902 | 0.925 | [0.822, 1.000] |
| `interpreter_advised_llm` | 115/120 | 0.958 | 0.656 [0.624, 0.691] | 0.667 | [0.443, 1.000] | 0.919 | 0.961 | [0.844, 1.000] |
| `all_interpreter_advised_llm` | 118/120 | 0.983 | 0.652 [0.621, 0.686] | 0.667 | [0.414, 1.000] | 0.914 | 0.961 | [0.840, 1.000] |
| `one_oracle` | 110/120 | 0.917 | 0.649 [0.612, 0.689] | 0.687 | [0.341, 1.000] | 0.904 | 0.961 | [0.829, 1.000] |
| `all_oracle` | 120/120 | 1.000 | 0.645 [0.612, 0.682] | 0.684 | [0.310, 1.000] | 0.907 | 0.982 | [0.843, 1.000] |

## Unconditional companions

`below-thr. episode` counts a no-deal episode as 0 rather than dropping it, so it is the rate at which an arm's episodes END in an individually irrational agreement.

| arm | n | deal rate | below-thr. episode | normalized score | norm. Nash welf. (uncond.) |
|---|---:|---:|---:|---:|---:|
| `all_llm` | 120 | 0.958 [0.925, 0.992] | 0.000 [0.000, 0.000] | 0.873 [0.833, 0.910] | 0.512 [0.429, 0.585] |
| `all_rational` | 120 | 0.233 [0.133, 0.342] | 0.000 [0.000, 0.000] | 0.189 [0.110, 0.277] | 0.095 [0.049, 0.145] |
| `one_rational` | 120 | 0.767 [0.683, 0.842] | 0.000 [0.000, 0.000] | 0.686 [0.614, 0.758] | 0.388 [0.319, 0.451] |
| `rational_advised_llm` | 120 | 0.900 [0.808, 0.975] | 0.000 [0.000, 0.000] | 0.811 [0.726, 0.884] | 0.473 [0.393, 0.546] |
| `interpreter_advised_llm` | 120 | 0.958 [0.917, 0.992] | 0.000 [0.000, 0.000] | 0.881 [0.836, 0.921] | 0.520 [0.441, 0.587] |
| `all_interpreter_advised_llm` | 120 | 0.983 [0.950, 1.000] | 0.000 [0.000, 0.000] | 0.899 [0.856, 0.935] | 0.502 [0.420, 0.574] |
| `one_oracle` | 120 | 0.917 [0.850, 0.967] | 0.000 [0.000, 0.000] | 0.829 [0.765, 0.885] | 0.433 [0.347, 0.516] |
| `all_oracle` | 120 | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | 0.907 [0.880, 0.933] | 0.406 [0.317, 0.495] |

## Paired contrasts by difficulty tag vs `all_rational`

Deal rate and normalized score, paired on `(instance, seed)` within each tag. A tag is a property of the parameter set, so tags overlap and the rows do not partition the bank.

| arm | metric | easy | hard | high-conflict | large-frontier | medium | no-clear-win-win | pivotal-seat | small-frontier |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `all_llm` | deal rate | 0.450 [0.275, 0.600] | 0.925 [0.825, 1.000] | 0.867 [0.733, 0.967] | 0.667 [0.433, 0.900] | 0.800 [0.650, 0.925] | 0.867 [0.733, 0.967] | 0.900 [0.767, 1.000] | 0.800 [0.600, 0.967] |
| `all_llm` | norm. score | 0.447 [0.327, 0.570] | 0.869 [0.771, 0.957] | 0.818 [0.690, 0.927] | 0.647 [0.456, 0.823] | 0.735 [0.544, 0.889] | 0.824 [0.708, 0.942] | 0.849 [0.735, 0.959] | 0.738 [0.475, 0.940] |
| `one_rational` | deal rate | 0.350 [0.225, 0.450] | 0.575 [0.450, 0.700] | 0.667 [0.533, 0.767] | 0.500 [0.333, 0.667] | 0.675 [0.500, 0.825] | 0.567 [0.400, 0.733] | 0.533 [0.367, 0.667] | 0.667 [0.467, 0.833] |
| `one_rational` | norm. score | 0.351 [0.242, 0.458] | 0.527 [0.416, 0.625] | 0.598 [0.478, 0.698] | 0.498 [0.336, 0.653] | 0.615 [0.408, 0.792] | 0.519 [0.371, 0.664] | 0.494 [0.366, 0.628] | 0.608 [0.360, 0.826] |
| `rational_advised_llm` | deal rate | 0.475 [0.325, 0.625] | 0.750 [0.500, 0.950] | 0.667 [0.433, 0.900] | 0.667 [0.533, 0.833] | 0.775 [0.649, 0.900] | 0.600 [0.333, 0.867] | 0.800 [0.533, 1.000] | 0.667 [0.433, 0.900] |
| `rational_advised_llm` | norm. score | 0.463 [0.339, 0.592] | 0.693 [0.474, 0.891] | 0.623 [0.409, 0.828] | 0.643 [0.538, 0.747] | 0.711 [0.515, 0.862] | 0.552 [0.329, 0.775] | 0.748 [0.508, 0.947] | 0.595 [0.320, 0.867] |
| `interpreter_advised_llm` | deal rate | 0.500 [0.325, 0.650] | 0.875 [0.725, 0.975] | 0.833 [0.733, 0.933] | 0.700 [0.533, 0.834] | 0.800 [0.675, 0.900] | 0.800 [0.633, 0.933] | 0.867 [0.667, 1.000] | 0.800 [0.633, 0.933] |
| `interpreter_advised_llm` | norm. score | 0.510 [0.376, 0.637] | 0.827 [0.690, 0.944] | 0.795 [0.669, 0.921] | 0.688 [0.563, 0.813] | 0.739 [0.572, 0.869] | 0.765 [0.609, 0.906] | 0.823 [0.653, 0.955] | 0.754 [0.538, 0.919] |
| `all_interpreter_advised_llm` | deal rate | 0.500 [0.325, 0.650] | 0.925 [0.775, 1.000] | 0.867 [0.733, 0.967] | 0.733 [0.567, 0.900] | 0.825 [0.675, 0.925] | 0.867 [0.667, 1.000] | 0.900 [0.700, 1.000] | 0.867 [0.667, 1.000] |
| `all_interpreter_advised_llm` | norm. score | 0.500 [0.379, 0.614] | 0.881 [0.745, 0.974] | 0.825 [0.712, 0.934] | 0.695 [0.554, 0.829] | 0.750 [0.543, 0.901] | 0.828 [0.656, 0.951] | 0.870 [0.690, 0.974] | 0.802 [0.524, 0.981] |
| `one_oracle` | deal rate | 0.450 [0.300, 0.575] | 0.850 [0.650, 1.000] | 0.833 [0.700, 0.967] | 0.667 [0.500, 0.833] | 0.750 [0.625, 0.850] | 0.800 [0.567, 0.967] | 0.800 [0.533, 0.967] | 0.800 [0.633, 0.933] |
| `one_oracle` | norm. score | 0.443 [0.335, 0.536] | 0.788 [0.606, 0.935] | 0.782 [0.634, 0.918] | 0.640 [0.478, 0.780] | 0.688 [0.535, 0.821] | 0.764 [0.540, 0.943] | 0.741 [0.502, 0.921] | 0.751 [0.540, 0.924] |
| `all_oracle` | deal rate | 0.500 [0.325, 0.650] | 0.975 [0.925, 1.000] | 0.867 [0.733, 0.967] | 0.733 [0.567, 0.900] | 0.825 [0.675, 0.925] | 0.933 [0.867, 1.000] | 0.967 [0.900, 1.000] | 0.867 [0.667, 1.000] |
| `all_oracle` | norm. score | 0.463 [0.343, 0.570] | 0.934 [0.888, 0.974] | 0.788 [0.632, 0.921] | 0.689 [0.537, 0.845] | 0.757 [0.583, 0.885] | 0.906 [0.813, 0.987] | 0.925 [0.869, 0.974] | 0.827 [0.585, 0.993] |


## Paired fairness contrasts vs `all_rational`

On the conditional columns a pair contributes only when BOTH arms closed a deal on that `(instance, seed)`; the pair count is reported because it falls sharply for the low-deal-rate arms.

| arm | deal rate | below-thr. accept | worst-off z | worst-off share | norm. Gini | max share |
|---|---:|---:|---:|---:|---:|---:|
| `all_llm` | 0.725 [0.608, 0.833] (n=120) | 0.000 [0.000, 0.000] (n=28) | 0.096 [0.023, 0.163] (n=28) | 0.028 [0.005, 0.050] (n=28) | -0.031 [-0.067, 0.004] (n=28) | -0.006 [-0.022, 0.013] (n=28) |
| `one_rational` | 0.533 [0.433, 0.625] (n=120) | 0.000 [0.000, 0.000] (n=26) | 0.122 [0.049, 0.172] (n=26) | 0.036 [0.011, 0.053] (n=26) | -0.053 [-0.086, -0.012] (n=26) | -0.019 [-0.039, 0.003] (n=26) |
| `rational_advised_llm` | 0.667 [0.550, 0.783] (n=120) | 0.000 [0.000, 0.000] (n=27) | 0.124 [0.052, 0.208] (n=27) | 0.036 [0.011, 0.063] (n=27) | -0.033 [-0.082, 0.008] (n=27) | -0.002 [-0.032, 0.028] (n=27) |
| `interpreter_advised_llm` | 0.725 [0.617, 0.825] (n=120) | 0.000 [0.000, 0.000] (n=28) | 0.137 [0.057, 0.212] (n=28) | 0.039 [0.014, 0.064] (n=28) | -0.046 [-0.087, -0.004] (n=28) | -0.012 [-0.032, 0.012] (n=28) |
| `all_interpreter_advised_llm` | 0.750 [0.633, 0.850] (n=120) | 0.000 [0.000, 0.000] (n=28) | 0.114 [0.033, 0.199] (n=28) | 0.034 [0.007, 0.062] (n=28) | -0.025 [-0.072, 0.019] (n=28) | -0.005 [-0.027, 0.017] (n=28) |
| `one_oracle` | 0.683 [0.567, 0.792] (n=120) | 0.000 [0.000, 0.000] (n=26) | 0.016 [-0.059, 0.081] (n=26) | 0.002 [-0.022, 0.024] (n=26) | 0.045 [0.006, 0.088] (n=26) | 0.028 [0.005, 0.056] (n=26) |
| `all_oracle` | 0.767 [0.658, 0.867] (n=120) | 0.000 [0.000, 0.000] (n=28) | 0.006 [-0.038, 0.047] (n=28) | -0.001 [-0.017, 0.014] (n=28) | 0.048 [0.017, 0.076] (n=28) | 0.031 [0.014, 0.050] (n=28) |

## Correct-walk credit

**Every one of the 24 bank instances admits at least one weakly all-IR deal** (u_i ≥ tau_i for all five parties; IR-set sizes 4–47 of 256 deals). The discrete justified-walk fraction is therefore identically 0.000 in every arm and every stratum: **no no-deal episode in this campaign was a correct refusal of an infeasible game.** The metric charging zero for a walk is not, on this bank, charging zero for correct behaviour.

The continuous form carries what is left. The IR margin of an instance is max over deals of min_i z_i — how much room the best all-satisfying deal leaves the party it satisfies least. Across the bank it ranges 0.000–0.652 (mean 0.377). But feasibility is not comfortable everywhere: 3 of 24 instances sit at or below a 0.10 margin, and 2 of those are exactly 0.000 — the best all-satisfying deal in those games holds some party precisely at its threshold, so agreement is feasible on paper and knife-edge in practice. Those instances carry the tags hard, high-conflict, no-clear-win-win, pivotal-seat, small-frontier, which is the same corner of the bank the inverted difficulty gradient points at.


### No-deal episodes by arm and difficulty tag

Each cell is `no-deals/episodes (justified walks)`, where a walk is justified only if its instance admits no all-IR deal. The walk-weighted mean IR margin of each cell is in `summary.json`.

| tag | `all_llm` | `all_rational` | `one_rational` | `rational_advised_llm` | `interpreter_advised_llm` | `all_interpreter_advised_llm` | `one_oracle` | `all_oracle` |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| easy | 2/40 (0 just.) | 20/40 (0 just.) | 6/40 (0 just.) | 1/40 (0 just.) | 0/40 (0 just.) | 0/40 (0 just.) | 2/40 (0 just.) | 0/40 (0 just.) |
| hard | 2/40 (0 just.) | 39/40 (0 just.) | 16/40 (0 just.) | 9/40 (0 just.) | 4/40 (0 just.) | 2/40 (0 just.) | 5/40 (0 just.) | 0/40 (0 just.) |
| high-conflict | 0/30 (0 just.) | 26/30 (0 just.) | 6/30 (0 just.) | 6/30 (0 just.) | 1/30 (0 just.) | 0/30 (0 just.) | 1/30 (0 just.) | 0/30 (0 just.) |
| large-frontier | 2/30 (0 just.) | 22/30 (0 just.) | 7/30 (0 just.) | 2/30 (0 just.) | 1/30 (0 just.) | 0/30 (0 just.) | 2/30 (0 just.) | 0/30 (0 just.) |
| medium | 1/40 (0 just.) | 33/40 (0 just.) | 6/40 (0 just.) | 2/40 (0 just.) | 1/40 (0 just.) | 0/40 (0 just.) | 3/40 (0 just.) | 0/40 (0 just.) |
| no-clear-win-win | 2/30 (0 just.) | 28/30 (0 just.) | 11/30 (0 just.) | 10/30 (0 just.) | 4/30 (0 just.) | 2/30 (0 just.) | 4/30 (0 just.) | 0/30 (0 just.) |
| pivotal-seat | 2/30 (0 just.) | 29/30 (0 just.) | 13/30 (0 just.) | 5/30 (0 just.) | 3/30 (0 just.) | 2/30 (0 just.) | 5/30 (0 just.) | 0/30 (0 just.) |
| small-frontier | 2/30 (0 just.) | 26/30 (0 just.) | 6/30 (0 just.) | 6/30 (0 just.) | 2/30 (0 just.) | 0/30 (0 just.) | 2/30 (0 just.) | 0/30 (0 just.) |

### Bank IR margin by difficulty tag

| tag | instances | mean IR margin | min IR margin |
|---|---:|---:|---:|
| easy | 8 | 0.442 | 0.329 |
| hard | 8 | 0.297 | 0.000 |
| high-conflict | 6 | 0.354 | 0.086 |
| large-frontier | 6 | 0.430 | 0.310 |
| medium | 8 | 0.394 | 0.269 |
| no-clear-win-win | 6 | 0.320 | 0.000 |
| pivotal-seat | 6 | 0.236 | 0.000 |
| small-frontier | 6 | 0.331 | 0.000 |

### Does no-deal track thin margins?

Per-arm Pearson correlation between an instance's IR margin and its no-deal rate in that arm (negative = failures concentrate where agreement was thinnest).

| arm | r(IR margin, no-deal rate) [95% CI] | instances |
|---|---:|---:|
| `all_llm` | -0.465 [-0.823, 0.097] | 24 |
| `all_rational` | -0.256 [-0.549, 0.079] | 24 |
| `one_rational` | -0.249 [-0.705, 0.316] | 24 |
| `rational_advised_llm` | -0.360 [-0.847, 0.172] | 24 |
| `interpreter_advised_llm` | -0.546 [-0.821, 0.104] | 24 |
| `all_interpreter_advised_llm` | -0.486 [-0.829, -0.368] | 24 |
| `one_oracle` | -0.506 [-0.781, 0.103] | 24 |
| `all_oracle` | — | 24 |
