Nine whole-threshold parameter sets × two matched seeds per arm. All LLM seats use Qwen3-8B with native thinking, datacenter framing, public persuasive chat, private score sheets, and unanimity. Each arm has 18/18 validated episodes.
The omniscient seat (OmniscientBestResponsePolicy) cast its forced-final vote on
whichever live offer it valued most instead of on the one offer under the up/down vote; the protocol rejected
that as a legality error, the seat spent its single retry repeating itself, and the turn was recorded as a
pass — a silent abstention. Fixed in commit ca20157, which postdates this campaign,
so every number for the oracle arm below is measured on a defective agent and its closure
figures in particular are not a negotiation result. The all-LLM and one-rational arms are unaffected.
This subset has not been re-run, so no corrected figure is offered here. For the size of the correction where it was measured: on the frozen Opus campaign, repairing the ballot moves the five-oracle arm's deal rate from 0.875 to 1.000, and on a fresh-bank replication the one-oracle paired score against all-LLM is −0.068 rather than −0.412. Full account in research notes 0045, 0039 and 0043; see the five-arm hub erratum.
Result so far: one rational seat was statistically indistinguishable from all-Qwen play (−0.036 score; interval crosses zero). One omniscient oracle seat improved score by +0.449, with the entire clustered interval above zero — but see the erratum above: that seat's ballot was spoiled, and this subset has not been re-run. This is an exploratory small subset, not the 120-episode Opus confirmatory campaign.
| Arm | Mean score [95% CI] | Deal rate [95% CI] | vs all LLM |
|---|---|---|---|
| All Qwen LLMs | 0.361 [0.137, 0.615] | 0.389 [0.167, 0.667] | reference |
| One rational | 0.325 [0.118, 0.531] | 0.333 [0.111, 0.556] | −0.036 [−0.364, 0.292] |
| One oracle (spoiled ballot — see erratum) | 0.810 [0.639, 0.941] | 0.833 [0.667, 1.000] | +0.449 [0.209, 0.692] |
Rotating omniscient oracle seat. Spoiled ballot — see the erratum above.
Open 18 visualizersIntervals use 10,000 bootstrap samples clustered by the nine parameter sets, keeping both seeds together. Every episode page includes turn-level rational and oracle counterfactual annotations.