Five-seat Opus rational-agent campaign
Complete private-information campaign: 24 whole-number parameter sets × five matched seeds × five arms = 600 episodes. LLM seats use Claude Opus 5 with thinking; all games use five-seat unanimity and the datacenter framing. Intervals resample parameter sets as clusters.
| Arm | Mean normalized score [95% CI] | Deals | Deal rate [95% CI] |
|---|---|---|---|
| Five Opus LLMs | 0.8726 [0.8328, 0.9104] | 115 / 120 | 0.9583 [0.9250, 0.9917] |
| Four Opus + one rational | 0.6864 [0.6132, 0.7594] | 92 / 120 | 0.7667 [0.6917, 0.8417] |
| Four Opus + one oracle (spoiled ballot — superseded) | 0.4607 [0.3771, 0.5436] | 61 / 120 | 0.5083 [0.4167, 0.6000] |
| Four Opus + one oracle — ballot repaired | 0.8290 [0.7669, 0.8838] | 110 / 120 | 0.9167 [0.8500, 0.9667] |
| Five rational agents | 0.1889 [0.1098, 0.2782] | 28 / 120 | 0.2333 [0.1333, 0.3417] |
| Five oracle agents (spoiled ballot — superseded) | 0.7907 [0.7343, 0.8417] | 105 / 120 | 0.8750 [0.8167, 0.9250] |
| Five oracle agents — ballot repaired | 0.9069 [0.8803, 0.9325] | 120 / 120 | 1.0000 [1.0000, 1.0000] |
Erratum (2026-08-10) — the oracle arms' closure numbers are a spoiled ballot.
The omniscient seat cast its forced-final vote on whichever live offer it valued most instead of the one
offer under the up/down vote; the protocol rejected that as a legality error, the seat repeated itself on
its single retry, and the turn was recorded as a pass. That consumed 94 of 107
(87.9%) of the five-oracle arm's forced-final turns — touching every one of its 15
no-deals — and 114 of 334 in the one-oracle arm (57 of its 59). The Opus and rational arms are
unaffected. Repaired and re-run on the identical games, the five-oracle arm closes
every game and its paired score against all-LLM flips from −0.082 to
+0.034 [+0.001, +0.074]. The distributional finding survives and sharpens: it now closes
more often than all-LLM (+0.042 [+0.017, +0.075]) and still splits worse on every column. Full
account in research note 0045 (repair commit
ca20157); details in
analysis.md and the
seeded-optimal-opening hub.
Erratum, part 2 (2026-08-14) — the one-oracle re-run landed, and it changes this campaign's
headline shape.
Repaired and re-run on the identical 120 games, the one-oracle arm goes from
0.508 to 0.917 deal rate and 0.461 to 0.829 score, and both of its paired
effects against all-LLM go to null: deal rate −0.450 → −0.042
[−0.100, +0.017], score −0.412 → −0.044 [−0.104, +0.011].
52 episodes flip from no-deal to deal and 3 flip the other way. What survives is the distributional cost,
unchanged in size and now on 106 paired closed games instead of 56 (worst-off z −0.060
[−0.092, −0.033], Gini +0.043 [+0.025, +0.065]).
With both oracle arms repaired the ordering is all-oracle > all-LLM > one-oracle >
one-rational > all-rational, so the published reading — computable agents collapse, worse the more
of them there are — does not survive: only the private-information Bayesian agents collapse, and the
gradient these five arms trace is information, not computability. Every computable arm still splits
worse than all-LLM at every composition. Recomputed basket over both repaired arms:
fairness_basket_ballot_repaired_both_oracles.md.
A repaired arm is a different agent from the frozen arm of the same name. The two rows are never pooled, averaged, or substituted for one another; contrasts are paired on the identical (instance, seed) games, bank, framing and protocol, but the arms are unpaired in time.
Whose ballots were dropped, split by seat occupant — a bare count across arms would attribute model errors to the harness and vice versa. Final-vote turns rejected for naming an offer other than the one under vote:
Every dropped ballot in the two oracle arms is the computable seat's; every one elsewhere is an LLM seat's,
at two orders of magnitude lower a rate and from a different cause (a model naming a wrong offer id). The
three non-oracle arms do not carry this defect.
A repaired arm is a different agent from the frozen arm of the same name. The two rows are never pooled, averaged, or substituted for one another; contrasts are paired on the identical (instance, seed) games, bank, framing and protocol, but the arms are unpaired in time.
Whose ballots were dropped, split by seat occupant — a bare count across arms would attribute model errors to the harness and vice versa. Final-vote turns rejected for naming an offer other than the one under vote:
| arm | computable seat | LLM seats |
|---|---|---|
| all-LLM | — | 3 / 31 |
| one-rational | 0 / 53 | 1 / 206 |
| one-oracle | 114 / 116 | 0 / 218 |
| all-rational | 0 / 476 | — |
| all-oracle | 94 / 107 | — |
Headline comparison
Five rational minus one rational: −0.4976 normalized score, 95% CI [−0.5901, −0.4036], and −53.3 percentage points of deal rate, 95% CI [−62.5, −43.3]. One rational agent was much less brittle than an entirely rational table, but five Opus agents still performed best.
Open episode visualizers
Five Opus LLMs — 120 episodes
Matched all-LLM vs one-rational — 120 pairs
One omniscient oracle — 120 episodes
Five rational agents — 120 episodes
Five omniscient oracles — 120 episodes
Every episode view includes the actual action plus the private-rational and omniscient-oracle turn counterfactuals.
