Five-seat Opus rational-agent campaign
Complete private-information campaign: 24 whole-number parameter sets × five matched seeds × five arms = 600 episodes. LLM seats use Claude Opus 5 with thinking; all games use five-seat unanimity and the datacenter framing. Intervals resample parameter sets as clusters.
| Arm | Mean normalized score [95% CI] | Deals | Deal rate [95% CI] |
|---|---|---|---|
| Five Opus LLMs | 0.8726 [0.8328, 0.9104] | 115 / 120 | 0.9583 [0.9250, 0.9917] |
| Four Opus + one rational | 0.6864 [0.6132, 0.7594] | 92 / 120 | 0.7667 [0.6917, 0.8417] |
| Four Opus + one oracle (spoiled ballot — re-run in flight) | 0.4607 [0.3771, 0.5436] | 61 / 120 | 0.5083 [0.4167, 0.6000] |
| Five rational agents | 0.1889 [0.1098, 0.2782] | 28 / 120 | 0.2333 [0.1333, 0.3417] |
| Five oracle agents (spoiled ballot — superseded) | 0.7907 [0.7343, 0.8417] | 105 / 120 | 0.8750 [0.8167, 0.9250] |
| Five oracle agents — ballot repaired | 0.9069 [0.8803, 0.9325] | 120 / 120 | 1.0000 [1.0000, 1.0000] |
Erratum (2026-08-10) — the oracle arms' closure numbers are a spoiled ballot.
The omniscient seat cast its forced-final vote on whichever live offer it valued most instead of the one
offer under the up/down vote; the protocol rejected that as a legality error, the seat repeated itself on
its single retry, and the turn was recorded as a pass. That consumed 94 of 107
(87.9%) of the five-oracle arm's forced-final turns — touching every one of its 15
no-deals — and 114 of 334 in the one-oracle arm (57 of its 59). The Opus and rational arms are
unaffected. Repaired and re-run on the identical games, the five-oracle arm closes
every game and its paired score against all-LLM flips from −0.082 to
+0.034 [+0.001, +0.074]. The distributional finding survives and sharpens: it now closes
more often than all-LLM (+0.042 [+0.017, +0.075]) and still splits worse on every column. Full
account in research note 0045 (repair commit
ca20157); details in
analysis.md and the
seeded-optimal-opening hub. The one-oracle re-run needs
API budget and is in flight.
Headline comparison
Five rational minus one rational: −0.4976 normalized score, 95% CI [−0.5901, −0.4036], and −53.3 percentage points of deal rate, 95% CI [−62.5, −43.3]. One rational agent was much less brittle than an entirely rational table, but five Opus agents still performed best.
Open episode visualizers
Five Opus LLMs — 120 episodes
Matched all-LLM vs one-rational — 120 pairs
One omniscient oracle — 120 episodes
Five rational agents — 120 episodes
Five omniscient oracles — 120 episodes
Every episode view includes the actual action plus the private-rational and omniscient-oracle turn counterfactuals.
