Does wanting the table's welfare change the table?
The five-arm campaign replaced frontier LLM negotiators with agents that are better at getting their own way. This extension seats agents that are exactly as capable and want something else: the same composed belief / optimal-stopping / best-response negotiator with its objective column swapped from own surplus to the table's normalized Nash welfare. Same 24 frozen games, same seeds, same private information, same Opus seats, paired episode by episode.
OmniscientBestResponsePolicy arms on this page carry a spoiled ballot, and the headline objective-swap number is inflated.
That agent cast its forced-final vote on whichever live offer it valued most instead of on the
one offer under the up/down vote; the protocol rejected that as a legality error, the seat spent its single
retry repeating itself, and the turn was recorded as a pass — a silent abstention. It consumed
114 of 334 forced-final turns in one_oracle and 94 of 107 in
all_oracle. Every other arm here is clean: all_llm, one_rational,
all_rational, all_selfish_dp_oracle and all four fairness arms re-derive with
0 mismatches over 9,589 turns (the fairness and DP seats are a different, composed policy).
Fixed in commit ca20157, which postdates this campaign.
What the repaired numbers say. On a preregistered fresh-bank replication run twice — once
against the spoiled agent, once against the repaired one on identical games —
one_oracle − all_llm on utilitarian score is
−0.068 [−0.136, −0.001], not −0.412, and its deal rate sits about
8 points below all-LLM rather than 45 (0.558 → 0.892 repaired).
all_oracle closes every game once repaired (0.875 → 1.000), so all 15 of its
no-deals were spoiled ballots. The +0.389 objective-swap contrast below is anchored on the
spoiled arm and is inflated by the same defect.
The “information amplifies motive” claim is withdrawn. Its interaction is
+0.019 [−0.052, +0.091] on the repaired primary — it contains zero, against the
+0.378 originally implied. What survives is a quarter-size interaction on the welfare coordinate (normalized
Nash welfare +0.071 at one seat, +0.137 at five) and the reading that information's first-order effect
is agreement, and it is motive-neutral; motive decides distribution. The all_oracle
distributional finding survives the repair and sharpens: repaired, it closes more often than all-LLM
(+0.042 [+0.017, +0.075]) and still splits worse on every column.
Nothing on this page has been deleted — the affected rows are marked (spoiled ballot) and kept as the record of what was computed. Do not quote them unqualified, and never pool spoiled with repaired cells. Full account: research notes 0039 (replication verdict), 0045 (the gate that found it) and 0043 (the fairness-basket erratum); details in analysis.md and on the five-arm hub.
Both co-primaries, side by side
Each family of agents optimizes a different quantity, and each leads on the one it optimizes —
so neither column is reported alone. utilitarian score is what the self-interested agents
maximize; normalized Nash welfare is the fairness objective's own coordinate. Deal rate and the
distributional columns are the referees. Intervals resample the 24 parameter sets. Extension arms are marked
with a rule on the left; all_selfish_dp_oracle is a control, not a result — it is the same
full-information composed DP with the own-surplus objective, which is what makes the fairness
contrast one-variable rather than a comparison across two different decision surfaces.
| arm | utilitarian score | normalized Nash welfare | deal rate | dist. to NBS | Gini | worst-off share | max share | no deal: walked | no deal: expired | turns to close | distinct openers |
|---|---|---|---|---|---|---|---|---|---|---|---|
all_llmfive Opus seats | 0.873[0.832, 0.910] | 0.512[0.430, 0.587] | 0.958[0.925, 0.992] | 0.464[0.359, 0.588] | 0.299[0.275, 0.328] | 0.257[0.206, 0.307] | 0.299[0.287, 0.309] | 0.000[0.000, 0.000] | 0.042[0.008, 0.075] | 17.748[17.079, 18.381] | 3.967[3.767, 4.167] |
one_rational4 Opus + 1 private self-interested Bayesian seat | 0.686[0.613, 0.760] | 0.388[0.318, 0.452] | 0.767[0.683, 0.842] | 0.444[0.344, 0.565] | 0.294[0.267, 0.325] | 0.239[0.195, 0.283] | 0.301[0.289, 0.314] | 0.000[0.000, 0.000] | 0.233[0.158, 0.317] | 21.391[20.623, 22.140] | 3.917[3.725, 4.108] |
one_oracle4 Opus + 1 omniscient self-interested seat (enumerate-and-score) — spoiled ballot; repaired paired score −0.068, see erratum | 0.461[0.377, 0.543] | 0.240[0.173, 0.308] | 0.508[0.417, 0.600] | 0.549[0.418, 0.687] | 0.346[0.319, 0.379] | 0.187[0.131, 0.246] | 0.314[0.298, 0.332] | 0.000[0.000, 0.000] | 0.492[0.400, 0.583] | 19.475[18.588, 20.453] | 3.883[3.708, 4.050] |
all_rationalfive private self-interested Bayesian seats | 0.189[0.110, 0.279] | 0.095[0.051, 0.146] | 0.233[0.133, 0.350] | 0.600[0.497, 0.680] | 0.311[0.284, 0.339] | 0.147[0.102, 0.197] | 0.314[0.299, 0.326] | 0.000[0.000, 0.000] | 0.767[0.650, 0.867] | 24.821[24.348, 25.000] | 3.600[3.375, 3.817] |
all_oraclefive omniscient self-interested seats (enumerate-and-score) — spoiled ballot; repaired deal rate 1.000, see erratum | 0.791[0.734, 0.842] | 0.370[0.285, 0.449] | 0.875[0.817, 0.925] | 0.621[0.524, 0.738] | 0.354[0.329, 0.384] | 0.172[0.127, 0.217] | 0.319[0.305, 0.334] | 0.000[0.000, 0.000] | 0.125[0.075, 0.183] | 5.000[5.000, 5.000] | 1.142[1.083, 1.208] |
all_selfish_dp_oraclefive omniscient self-interested seats (composed DP) — the decision-surface control | 0.871[0.841, 0.902] | 0.397[0.303, 0.489] | 1.000[1.000, 1.000] | 0.609[0.476, 0.752] | 0.356[0.320, 0.394] | 0.161[0.109, 0.217] | 0.329[0.311, 0.349] | 0.000[0.000, 0.000] | 0.000[0.000, 0.000] | 24.783[24.508, 25.000] | 3.175[2.883, 3.467] |
one_fairness_oracle4 Opus + 1 omniscient fairness seat | 0.850[0.809, 0.890] | 0.492[0.412, 0.562] | 0.933[0.892, 0.967] | 0.391[0.283, 0.524] | 0.296[0.265, 0.331] | 0.273[0.222, 0.323] | 0.299[0.288, 0.309] | 0.000[0.000, 0.000] | 0.067[0.033, 0.108] | 18.964[18.434, 19.505] | 4.025[3.883, 4.175] |
one_fairness_algorithmic4 Opus + 1 private-information fairness seat | 0.697[0.617, 0.773] | 0.408[0.338, 0.472] | 0.808[0.725, 0.883] | 0.550[0.425, 0.708] | 0.312[0.280, 0.352] | 0.233[0.190, 0.273] | 0.309[0.296, 0.322] | 0.000[0.000, 0.000] | 0.192[0.117, 0.275] | 19.732[18.927, 20.563] | 3.658[3.492, 3.808] |
all_fairness_oraclefive omniscient fairness seats | 0.967[0.945, 0.985] | 0.596[0.508, 0.667] | 1.000[1.000, 1.000] | 0.091[0.000, 0.249] | 0.252[0.212, 0.299] | 0.357[0.291, 0.416] | 0.275[0.265, 0.286] | 0.000[0.000, 0.000] | 0.000[0.000, 0.000] | 5.000[5.000, 5.000] | 1.000[1.000, 1.000] |
all_fairness_algorithmicfive private-information fairness seats | 0.278[0.151, 0.416] | 0.152[0.084, 0.227] | 0.375[0.217, 0.542] | 0.567[0.398, 0.732] | 0.324[0.304, 0.346] | 0.129[0.098, 0.161] | 0.334[0.314, 0.355] | 0.000[0.000, 0.000] | 0.625[0.458, 0.783] | 13.000[10.306, 16.081] | 2.300[2.058, 2.533] |
Paired effects vs all_llm
Every contrast joins on (instance_id, episode_seed), so arms are compared on
identical games. No-deal scores zero.
| arm | utilitarian score | normalized Nash welfare | deal rate | Gini | worst-off share | max share |
|---|---|---|---|---|---|---|
one_rational4 Opus + 1 private self-interested Bayesian seat | -0.186[-0.266, -0.106] | -0.124[-0.187, -0.063] | -0.192[-0.275, -0.108] | -0.005[-0.022, 0.011] | -0.011[-0.043, 0.021] | 0.000[-0.009, 0.010] |
one_oracle4 Opus + 1 omniscient self-interested seat (enumerate-and-score) — spoiled ballot; repaired paired score −0.068, see erratum | -0.412[-0.494, -0.323] | -0.273[-0.344, -0.198] | -0.450[-0.542, -0.350] | 0.040[0.014, 0.066] | -0.070[-0.114, -0.028] | 0.018[0.004, 0.034] |
all_rationalfive private self-interested Bayesian seats | -0.684[-0.786, -0.575] | -0.417[-0.509, -0.324] | -0.725[-0.833, -0.608] | 0.030[-0.003, 0.057] | -0.096[-0.163, -0.023] | 0.006[-0.013, 0.022] |
all_oraclefive omniscient self-interested seats (enumerate-and-score) — spoiled ballot; repaired deal rate 1.000, see erratum | -0.082[-0.145, -0.021] | -0.143[-0.208, -0.081] | -0.083[-0.150, -0.017] | 0.053[0.032, 0.074] | -0.075[-0.118, -0.032] | 0.019[0.006, 0.032] |
all_selfish_dp_oraclefive omniscient self-interested seats (composed DP) — the decision-surface control | -0.001[-0.041, 0.045] | -0.116[-0.180, -0.057] | 0.042[0.008, 0.075] | 0.054[0.023, 0.083] | -0.090[-0.133, -0.046] | 0.031[0.015, 0.048] |
one_fairness_oracle4 Opus + 1 omniscient fairness seat | -0.023[-0.070, 0.027] | -0.020[-0.055, 0.016] | -0.025[-0.067, 0.017] | -0.004[-0.021, 0.013] | 0.018[-0.006, 0.043] | -0.001[-0.007, 0.004] |
one_fairness_algorithmic4 Opus + 1 private-information fairness seat | -0.175[-0.258, -0.094] | -0.104[-0.161, -0.049] | -0.150[-0.233, -0.067] | 0.006[-0.014, 0.028] | -0.021[-0.049, 0.006] | 0.008[-0.002, 0.018] |
all_fairness_oraclefive omniscient fairness seats | 0.094[0.053, 0.139] | 0.083[0.057, 0.111] | 0.042[0.008, 0.075] | -0.053[-0.077, -0.027] | 0.108[0.072, 0.145] | -0.024[-0.036, -0.014] |
all_fairness_algorithmicfive private-information fairness seats | -0.595[-0.734, -0.448] | -0.361[-0.462, -0.256] | -0.583[-0.742, -0.417] | 0.058[0.035, 0.082] | -0.134[-0.204, -0.069] | 0.032[0.005, 0.059] |
Per-arm episode pages
all_fairness_algorithmic— every seat is a computable policy, so the episodes carry no model text and the tables above summarize them; pages on the cluster at/juice2/scr2/siddharth/ii_mats/rational_agents/fivepriv_opus_fairness_v2__all_fairness_algorithmic/visualizations/index.htmlall_fairness_oracle— every seat is a computable policy, so the episodes carry no model text and the tables above summarize them; pages on the cluster at/juice2/scr2/siddharth/ii_mats/rational_agents/fivepriv_opus_fairness_v2__all_fairness_oracle/visualizations/index.htmlall_selfish_dp_oracle— every seat is a computable policy, so the episodes carry no model text and the tables above summarize them; pages on the cluster at/juice2/scr2/siddharth/ii_mats/rational_agents/fivepriv_opus_fairness_v2__all_selfish_dp_oracle/visualizations/index.htmlone_fairness_algorithmic— per-episode pages (transcript, oracle counterfactuals, Pareto geometry)one_fairness_oracle— per-episode pages (transcript, oracle counterfactuals, Pareto geometry)
Artifacts
- analysis.md — the full rendered table set
- summary.json — every estimate and interval
- episode_rows.csv — one row per episode, all ten arms
- the frozen five-arm hub — unchanged; this page is a sibling