Completed exploratory subset: nine private-information parameter sets × seeds 0 and 1. Every new abstract arm has 18/18 valid episodes. Final clustered intervals and Qwen similarity are pending the unified robustness analysis.
The omniscient seat (OmniscientBestResponsePolicy) cast its forced-final vote on
whichever live offer it valued most instead of on the one offer under the up/down vote; the protocol rejected
that as a legality error, the seat spent its single retry repeating itself, and the turn was recorded as a
pass — a silent abstention. Fixed in commit ca20157, which postdates this campaign,
so every number for the oracle arm below is measured on a defective agent and its closure
figures in particular are not a negotiation result. The all-LLM and one-rational arms are unaffected.
This subset has not been re-run, so no corrected figure is offered here. For the size of the correction where it was measured: on the frozen Opus campaign, repairing the ballot moves the five-oracle arm's deal rate from 0.875 to 1.000, and on a fresh-bank replication the one-oracle paired score against all-LLM is −0.068 rather than −0.412. Full account in research notes 0045, 0039 and 0043; see the five-arm hub erratum.
| Arm | Abstract score | Datacenter score | Difference | Deals |
|---|---|---|---|---|
| All LLM | 0.93857 | 0.93009 | +0.00848 | 18 / 18 |
| One rational | 0.71533 | 0.83881 | −0.12348 | 15 / 17 |
| One oracle (spoiled ballot — see erratum) | 0.35694 | 0.41858 | −0.06164 | 7 / 8 |