Claude Opus 5: datacenter vs abstract framing

Completed exploratory subset: nine private-information parameter sets × seeds 0 and 1. Every new abstract arm has 18/18 valid episodes. Final clustered intervals and Qwen similarity are pending the unified robustness analysis.

⚠ Erratum (2026-08-10) — the one oracle arm on this page carries a spoiled ballot.

The omniscient seat (OmniscientBestResponsePolicy) cast its forced-final vote on whichever live offer it valued most instead of on the one offer under the up/down vote; the protocol rejected that as a legality error, the seat spent its single retry repeating itself, and the turn was recorded as a pass — a silent abstention. Fixed in commit ca20157, which postdates this campaign, so every number for the oracle arm below is measured on a defective agent and its closure figures in particular are not a negotiation result. The all-LLM and one-rational arms are unaffected.

This subset has not been re-run, so no corrected figure is offered here. For the size of the correction where it was measured: on the frozen Opus campaign, repairing the ballot moves the five-oracle arm's deal rate from 0.875 to 1.000, and on a fresh-bank replication the one-oracle paired score against all-LLM is −0.068 rather than −0.412. Full account in research notes 0045, 0039 and 0043; see the five-arm hub erratum.

ArmAbstract scoreDatacenter scoreDifferenceDeals
All LLM0.938570.93009+0.0084818 / 18
One rational0.715330.83881−0.1234815 / 17
One oracle (spoiled ballot — see erratum)0.356940.41858−0.061647 / 8