# Five-arm fairness basket and correct-walk credit

All intervals are 95% cluster bootstraps (10000 resamples) over the 24 parameter sets, each set's five seeds resampled together. Paired contrasts join on `(instance_id, episode_seed)`.


## Fairness basket — conditional on a closed deal

Every column below is computed over closed deals only, so `deal rate` is carried beside them: an arm that closes 28 of 120 deals is describing a self-selected set of games, not the same games. Normalized coordinate z_i = (u_i − tau_i)/c_i with c_i the party's best surplus on the individually rational set (affine-invariant per party).

| arm | n | deal rate | below-thr. accept | worst-off z | worst-off share | dist-NBS | dist-KS | norm. Gini | norm. Nash welf. | max share |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| `all_llm` | 120 | 0.958 [0.925, 0.992] | 0.000 [0.000, 0.000] | 0.257 [0.206, 0.306] | 0.077 [0.064, 0.090] | 0.464 [0.357, 0.586] | 0.398 [0.316, 0.474] | 0.223 [0.200, 0.247] | 0.535 [0.458, 0.601] | 0.299 [0.287, 0.309] |
| `all_rational` | 120 | 0.233 [0.133, 0.342] | 0.000 [0.000, 0.000] | 0.147 [0.103, 0.196] | 0.052 [0.037, 0.068] | 0.600 [0.498, 0.681] | 0.590 [0.520, 0.649] | 0.261 [0.238, 0.283] | 0.408 [0.318, 0.486] | 0.314 [0.299, 0.327] |
| `one_rational` | 120 | 0.767 [0.683, 0.842] | 0.000 [0.000, 0.000] | 0.239 [0.195, 0.282] | 0.073 [0.061, 0.083] | 0.444 [0.341, 0.564] | 0.414 [0.332, 0.493] | 0.231 [0.209, 0.255] | 0.506 [0.439, 0.565] | 0.301 [0.289, 0.314] |
| `rational_advised_llm` | 120 | 0.900 [0.808, 0.975] | 0.000 [0.000, 0.000] | 0.250 [0.206, 0.294] | 0.076 [0.064, 0.087] | 0.478 [0.385, 0.581] | 0.458 [0.372, 0.535] | 0.232 [0.209, 0.253] | 0.526 [0.465, 0.581] | 0.302 [0.291, 0.313] |
| `interpreter_advised_llm` | 120 | 0.958 [0.917, 0.992] | 0.000 [0.000, 0.000] | 0.269 [0.224, 0.310] | 0.082 [0.069, 0.093] | 0.412 [0.311, 0.536] | 0.366 [0.295, 0.430] | 0.214 [0.196, 0.233] | 0.543 [0.477, 0.598] | 0.296 [0.285, 0.307] |
| `all_interpreter_advised_llm` | 120 | 0.983 [0.950, 1.000] | 0.000 [0.000, 0.000] | 0.239 [0.193, 0.286] | 0.073 [0.060, 0.085] | 0.500 [0.392, 0.633] | 0.442 [0.353, 0.524] | 0.229 [0.207, 0.250] | 0.511 [0.436, 0.576] | 0.300 [0.288, 0.311] |
| `one_oracle` | 120 | 0.917 [0.850, 0.967] | 0.000 [0.000, 0.000] | 0.201 [0.151, 0.252] | 0.059 [0.047, 0.070] | 0.550 [0.436, 0.676] | 0.524 [0.426, 0.614] | 0.266 [0.237, 0.294] | 0.472 [0.393, 0.548] | 0.314 [0.297, 0.332] |
| `all_oracle` | 120 | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | 0.166 [0.118, 0.220] | 0.048 [0.037, 0.061] | 0.631 [0.523, 0.758] | 0.549 [0.450, 0.648] | 0.285 [0.255, 0.314] | 0.406 [0.317, 0.495] | 0.319 [0.303, 0.336] |

Lower is better for below-threshold accept, dist-NBS, dist-KS, Gini, and max share; higher is better for worst-off z, worst-off share, and normalized Nash welfare. An equal split puts worst-off share at 0.200 and max share at 0.200.


### Among-deals surplus (per-party z and episode score)

Both blocks condition on a closed deal, so deal rate is carried beside them: an arm's closed set is self-selected, and a low-deal-rate arm's columns describe a different, easier subset of games. Per-party z pools party-observations (five per closed deal, no per-deal averaging; mean with an instance-clustered bootstrap CI, quartiles empirical/descriptive); the episode-score block is descriptive (the unconditional headline already carries the arm's interval).

| arm | deals | deal rate | per-party z mean [95% CI] | z median | z IQR | score mean | score median | score IQR |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `all_llm` | 115/120 | 0.958 | 0.653 [0.619, 0.691] | 0.667 | [0.433, 1.000] | 0.911 | 0.946 | [0.838, 1.000] |
| `all_rational` | 28/120 | 0.233 | 0.554 [0.512, 0.611] | 0.567 | [0.329, 0.762] | 0.809 | 0.830 | [0.699, 0.922] |
| `one_rational` | 92/120 | 0.767 | 0.639 [0.606, 0.675] | 0.657 | [0.431, 1.000] | 0.895 | 0.931 | [0.825, 1.000] |
| `rational_advised_llm` | 108/120 | 0.900 | 0.646 [0.616, 0.678] | 0.667 | [0.414, 1.000] | 0.902 | 0.925 | [0.822, 1.000] |
| `interpreter_advised_llm` | 115/120 | 0.958 | 0.656 [0.624, 0.691] | 0.667 | [0.443, 1.000] | 0.919 | 0.961 | [0.844, 1.000] |
| `all_interpreter_advised_llm` | 118/120 | 0.983 | 0.652 [0.621, 0.686] | 0.667 | [0.414, 1.000] | 0.914 | 0.961 | [0.840, 1.000] |
| `one_oracle` | 110/120 | 0.917 | 0.649 [0.612, 0.689] | 0.687 | [0.341, 1.000] | 0.904 | 0.961 | [0.829, 1.000] |
| `all_oracle` | 120/120 | 1.000 | 0.645 [0.612, 0.682] | 0.684 | [0.310, 1.000] | 0.907 | 0.982 | [0.843, 1.000] |

## Unconditional companions

`below-thr. episode` counts a no-deal episode as 0 rather than dropping it, so it is the rate at which an arm's episodes END in an individually irrational agreement.

| arm | n | deal rate | below-thr. episode | normalized score | norm. Nash welf. (uncond.) |
|---|---:|---:|---:|---:|---:|
| `all_llm` | 120 | 0.958 [0.925, 0.992] | 0.000 [0.000, 0.000] | 0.873 [0.833, 0.910] | 0.512 [0.429, 0.585] |
| `all_rational` | 120 | 0.233 [0.133, 0.342] | 0.000 [0.000, 0.000] | 0.189 [0.110, 0.277] | 0.095 [0.049, 0.145] |
| `one_rational` | 120 | 0.767 [0.683, 0.842] | 0.000 [0.000, 0.000] | 0.686 [0.614, 0.758] | 0.388 [0.319, 0.451] |
| `rational_advised_llm` | 120 | 0.900 [0.808, 0.975] | 0.000 [0.000, 0.000] | 0.811 [0.726, 0.884] | 0.473 [0.393, 0.546] |
| `interpreter_advised_llm` | 120 | 0.958 [0.917, 0.992] | 0.000 [0.000, 0.000] | 0.881 [0.836, 0.921] | 0.520 [0.441, 0.587] |
| `all_interpreter_advised_llm` | 120 | 0.983 [0.950, 1.000] | 0.000 [0.000, 0.000] | 0.899 [0.856, 0.935] | 0.502 [0.420, 0.574] |
| `one_oracle` | 120 | 0.917 [0.850, 0.967] | 0.000 [0.000, 0.000] | 0.829 [0.765, 0.885] | 0.433 [0.347, 0.516] |
| `all_oracle` | 120 | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | 0.907 [0.880, 0.933] | 0.406 [0.317, 0.495] |

## Paired contrasts by difficulty tag vs `all_llm`

Deal rate and normalized score, paired on `(instance, seed)` within each tag. A tag is a property of the parameter set, so tags overlap and the rows do not partition the bank.

| arm | metric | easy | hard | high-conflict | large-frontier | medium | no-clear-win-win | pivotal-seat | small-frontier |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `all_rational` | deal rate | -0.450 [-0.600, -0.275] | -0.925 [-1.000, -0.825] | -0.867 [-0.967, -0.733] | -0.667 [-0.900, -0.433] | -0.800 [-0.925, -0.650] | -0.867 [-0.967, -0.733] | -0.900 [-1.000, -0.767] | -0.800 [-0.967, -0.600] |
| `all_rational` | norm. score | -0.447 [-0.570, -0.327] | -0.869 [-0.957, -0.771] | -0.818 [-0.927, -0.690] | -0.647 [-0.823, -0.456] | -0.735 [-0.889, -0.544] | -0.824 [-0.942, -0.708] | -0.849 [-0.959, -0.735] | -0.738 [-0.940, -0.475] |
| `one_rational` | deal rate | -0.100 [-0.225, 0.025] | -0.350 [-0.450, -0.250] | -0.200 [-0.333, -0.067] | -0.167 [-0.400, 0.033] | -0.125 [-0.300, 0.025] | -0.300 [-0.467, -0.166] | -0.367 [-0.467, -0.267] | -0.133 [-0.267, 0.033] |
| `one_rational` | norm. score | -0.096 [-0.197, 0.025] | -0.342 [-0.435, -0.259] | -0.220 [-0.327, -0.112] | -0.149 [-0.359, 0.066] | -0.120 [-0.276, 0.009] | -0.305 [-0.465, -0.166] | -0.356 [-0.466, -0.260] | -0.130 [-0.267, 0.022] |
| `rational_advised_llm` | deal rate | 0.025 [-0.050, 0.100] | -0.175 [-0.375, -0.025] | -0.200 [-0.467, -0.033] | 0.000 [-0.133, 0.133] | -0.025 [-0.100, 0.050] | -0.267 [-0.500, -0.067] | -0.100 [-0.233, 0.000] | -0.133 [-0.400, 0.067] |
| `rational_advised_llm` | norm. score | 0.016 [-0.059, 0.094] | -0.176 [-0.380, -0.017] | -0.194 [-0.458, -0.031] | -0.004 [-0.117, 0.110] | -0.024 [-0.099, 0.052] | -0.272 [-0.496, -0.088] | -0.101 [-0.229, 0.002] | -0.143 [-0.421, 0.058] |
| `interpreter_advised_llm` | deal rate | 0.050 [0.000, 0.125] | -0.050 [-0.125, 0.000] | -0.033 [-0.100, 0.000] | 0.033 [-0.067, 0.133] | 0.000 [-0.075, 0.075] | -0.067 [-0.133, 0.000] | -0.033 [-0.100, 0.000] | 0.000 [-0.100, 0.100] |
| `interpreter_advised_llm` | norm. score | 0.063 [0.009, 0.131] | -0.042 [-0.111, 0.019] | -0.022 [-0.088, 0.022] | 0.041 [-0.062, 0.152] | 0.004 [-0.069, 0.075] | -0.059 [-0.137, 0.015] | -0.027 [-0.106, 0.039] | 0.016 [-0.084, 0.107] |
| `all_interpreter_advised_llm` | deal rate | 0.050 [0.000, 0.125] | 0.000 [-0.075, 0.075] | 0.000 [0.000, 0.000] | 0.067 [0.000, 0.133] | 0.025 [0.000, 0.075] | 0.000 [-0.100, 0.100] | 0.000 [-0.100, 0.100] | 0.067 [0.000, 0.133] |
| `all_interpreter_advised_llm` | norm. score | 0.053 [0.008, 0.113] | 0.012 [-0.069, 0.101] | 0.008 [-0.015, 0.031] | 0.048 [-0.011, 0.128] | 0.014 [-0.025, 0.070] | 0.004 [-0.104, 0.128] | 0.021 [-0.089, 0.141] | 0.064 [-0.025, 0.167] |
| `one_oracle` | deal rate | 0.000 [-0.100, 0.100] | -0.075 [-0.175, 0.000] | -0.033 [-0.100, 0.000] | 0.000 [-0.133, 0.133] | -0.050 [-0.150, 0.050] | -0.067 [-0.200, 0.000] | -0.100 [-0.233, 0.000] | 0.000 [-0.100, 0.100] |
| `one_oracle` | norm. score | -0.003 [-0.098, 0.095] | -0.081 [-0.195, 0.004] | -0.036 [-0.124, 0.018] | -0.007 [-0.112, 0.111] | -0.047 [-0.137, 0.045] | -0.060 [-0.199, 0.015] | -0.108 [-0.246, 0.005] | 0.013 [-0.081, 0.096] |
| `all_oracle` | deal rate | 0.050 [0.000, 0.125] | 0.050 [0.000, 0.100] | 0.000 [0.000, 0.000] | 0.067 [0.000, 0.133] | 0.025 [0.000, 0.075] | 0.067 [0.000, 0.133] | 0.067 [0.000, 0.133] | 0.067 [0.000, 0.133] |
| `all_oracle` | norm. score | 0.016 [-0.028, 0.064] | 0.065 [-0.006, 0.155] | -0.030 [-0.065, 0.013] | 0.042 [-0.015, 0.094] | 0.022 [-0.025, 0.070] | 0.082 [-0.016, 0.197] | 0.076 [-0.017, 0.196] | 0.089 [0.006, 0.197] |


## Paired fairness contrasts vs `all_llm`

On the conditional columns a pair contributes only when BOTH arms closed a deal on that `(instance, seed)`; the pair count is reported because it falls sharply for the low-deal-rate arms.

| arm | deal rate | below-thr. accept | worst-off z | worst-off share | norm. Gini | max share |
|---|---:|---:|---:|---:|---:|---:|
| `all_rational` | -0.725 [-0.833, -0.608] (n=120) | 0.000 [0.000, 0.000] (n=28) | -0.096 [-0.163, -0.023] (n=28) | -0.028 [-0.050, -0.005] (n=28) | 0.031 [-0.004, 0.067] (n=28) | 0.006 [-0.013, 0.022] (n=28) |
| `one_rational` | -0.192 [-0.275, -0.108] (n=120) | 0.000 [0.000, 0.000] (n=90) | -0.011 [-0.043, 0.020] (n=90) | -0.004 [-0.015, 0.006] (n=90) | 0.005 [-0.016, 0.026] (n=90) | 0.000 [-0.009, 0.010] (n=90) |
| `rational_advised_llm` | -0.058 [-0.150, 0.017] (n=120) | 0.000 [0.000, 0.000] (n=105) | -0.005 [-0.027, 0.018] (n=105) | -0.002 [-0.009, 0.005] (n=105) | 0.009 [-0.005, 0.026] (n=105) | 0.003 [-0.004, 0.010] (n=105) |
| `interpreter_advised_llm` | 0.000 [-0.042, 0.042] (n=120) | 0.000 [0.000, 0.000] (n=112) | 0.011 [-0.012, 0.034] (n=112) | 0.004 [-0.002, 0.011] (n=112) | -0.009 [-0.020, 0.001] (n=112) | -0.003 [-0.009, 0.002] (n=112) |
| `all_interpreter_advised_llm` | 0.025 [-0.008, 0.058] (n=120) | 0.000 [0.000, 0.000] (n=113) | -0.015 [-0.044, 0.013] (n=113) | -0.004 [-0.012, 0.005] (n=113) | 0.005 [-0.007, 0.017] (n=113) | 0.001 [-0.004, 0.007] (n=113) |
| `one_oracle` | -0.042 [-0.100, 0.017] (n=120) | 0.000 [0.000, 0.000] (n=106) | -0.060 [-0.092, -0.032] (n=106) | -0.020 [-0.030, -0.011] (n=106) | 0.043 [0.025, 0.065] (n=106) | 0.015 [0.005, 0.028] (n=106) |
| `all_oracle` | 0.042 [0.008, 0.075] (n=120) | 0.000 [0.000, 0.000] (n=115) | -0.086 [-0.130, -0.043] (n=115) | -0.028 [-0.042, -0.015] (n=115) | 0.061 [0.033, 0.091] (n=115) | 0.020 [0.007, 0.035] (n=115) |

## Correct-walk credit

**Every one of the 24 bank instances admits at least one weakly all-IR deal** (u_i ≥ tau_i for all five parties; IR-set sizes 4–47 of 256 deals). The discrete justified-walk fraction is therefore identically 0.000 in every arm and every stratum: **no no-deal episode in this campaign was a correct refusal of an infeasible game.** The metric charging zero for a walk is not, on this bank, charging zero for correct behaviour.

The continuous form carries what is left. The IR margin of an instance is max over deals of min_i z_i — how much room the best all-satisfying deal leaves the party it satisfies least. Across the bank it ranges 0.000–0.652 (mean 0.377). But feasibility is not comfortable everywhere: 3 of 24 instances sit at or below a 0.10 margin, and 2 of those are exactly 0.000 — the best all-satisfying deal in those games holds some party precisely at its threshold, so agreement is feasible on paper and knife-edge in practice. Those instances carry the tags hard, high-conflict, no-clear-win-win, pivotal-seat, small-frontier, which is the same corner of the bank the inverted difficulty gradient points at.


### No-deal episodes by arm and difficulty tag

Each cell is `no-deals/episodes (justified walks)`, where a walk is justified only if its instance admits no all-IR deal. The walk-weighted mean IR margin of each cell is in `summary.json`.

| tag | `all_llm` | `all_rational` | `one_rational` | `rational_advised_llm` | `interpreter_advised_llm` | `all_interpreter_advised_llm` | `one_oracle` | `all_oracle` |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| easy | 2/40 (0 just.) | 20/40 (0 just.) | 6/40 (0 just.) | 1/40 (0 just.) | 0/40 (0 just.) | 0/40 (0 just.) | 2/40 (0 just.) | 0/40 (0 just.) |
| hard | 2/40 (0 just.) | 39/40 (0 just.) | 16/40 (0 just.) | 9/40 (0 just.) | 4/40 (0 just.) | 2/40 (0 just.) | 5/40 (0 just.) | 0/40 (0 just.) |
| high-conflict | 0/30 (0 just.) | 26/30 (0 just.) | 6/30 (0 just.) | 6/30 (0 just.) | 1/30 (0 just.) | 0/30 (0 just.) | 1/30 (0 just.) | 0/30 (0 just.) |
| large-frontier | 2/30 (0 just.) | 22/30 (0 just.) | 7/30 (0 just.) | 2/30 (0 just.) | 1/30 (0 just.) | 0/30 (0 just.) | 2/30 (0 just.) | 0/30 (0 just.) |
| medium | 1/40 (0 just.) | 33/40 (0 just.) | 6/40 (0 just.) | 2/40 (0 just.) | 1/40 (0 just.) | 0/40 (0 just.) | 3/40 (0 just.) | 0/40 (0 just.) |
| no-clear-win-win | 2/30 (0 just.) | 28/30 (0 just.) | 11/30 (0 just.) | 10/30 (0 just.) | 4/30 (0 just.) | 2/30 (0 just.) | 4/30 (0 just.) | 0/30 (0 just.) |
| pivotal-seat | 2/30 (0 just.) | 29/30 (0 just.) | 13/30 (0 just.) | 5/30 (0 just.) | 3/30 (0 just.) | 2/30 (0 just.) | 5/30 (0 just.) | 0/30 (0 just.) |
| small-frontier | 2/30 (0 just.) | 26/30 (0 just.) | 6/30 (0 just.) | 6/30 (0 just.) | 2/30 (0 just.) | 0/30 (0 just.) | 2/30 (0 just.) | 0/30 (0 just.) |

### Bank IR margin by difficulty tag

| tag | instances | mean IR margin | min IR margin |
|---|---:|---:|---:|
| easy | 8 | 0.442 | 0.329 |
| hard | 8 | 0.297 | 0.000 |
| high-conflict | 6 | 0.354 | 0.086 |
| large-frontier | 6 | 0.430 | 0.310 |
| medium | 8 | 0.394 | 0.269 |
| no-clear-win-win | 6 | 0.320 | 0.000 |
| pivotal-seat | 6 | 0.236 | 0.000 |
| small-frontier | 6 | 0.331 | 0.000 |

### Does no-deal track thin margins?

Per-arm Pearson correlation between an instance's IR margin and its no-deal rate in that arm (negative = failures concentrate where agreement was thinnest).

| arm | r(IR margin, no-deal rate) [95% CI] | instances |
|---|---:|---:|
| `all_llm` | -0.465 [-0.823, 0.097] | 24 |
| `all_rational` | -0.256 [-0.549, 0.079] | 24 |
| `one_rational` | -0.249 [-0.705, 0.316] | 24 |
| `rational_advised_llm` | -0.360 [-0.847, 0.172] | 24 |
| `interpreter_advised_llm` | -0.546 [-0.821, 0.104] | 24 |
| `all_interpreter_advised_llm` | -0.486 [-0.829, -0.368] | 24 |
| `one_oracle` | -0.506 [-0.781, 0.103] | 24 |
| `all_oracle` | — | 24 |
