Qwen3-8B private negotiation subset

Nine whole-threshold parameter sets × two matched seeds per arm. All LLM seats use Qwen3-8B with native thinking, datacenter framing, public persuasive chat, private score sheets, and unanimity. Each arm has 18/18 validated episodes.

⚠ Erratum (2026-08-10) — the one oracle arm on this page carries a spoiled ballot.

The omniscient seat (OmniscientBestResponsePolicy) cast its forced-final vote on whichever live offer it valued most instead of on the one offer under the up/down vote; the protocol rejected that as a legality error, the seat spent its single retry repeating itself, and the turn was recorded as a pass — a silent abstention. Fixed in commit ca20157, which postdates this campaign, so every number for the oracle arm below is measured on a defective agent and its closure figures in particular are not a negotiation result. The all-LLM and one-rational arms are unaffected.

This subset has not been re-run, so no corrected figure is offered here. For the size of the correction where it was measured: on the frozen Opus campaign, repairing the ballot moves the five-oracle arm's deal rate from 0.875 to 1.000, and on a fresh-bank replication the one-oracle paired score against all-LLM is −0.068 rather than −0.412. Full account in research notes 0045, 0039 and 0043; see the five-arm hub erratum.

Result so far: one rational seat was statistically indistinguishable from all-Qwen play (−0.036 score; interval crosses zero). One omniscient oracle seat improved score by +0.449, with the entire clustered interval above zero — but see the erratum above: that seat's ballot was spoiled, and this subset has not been re-run. This is an exploratory small subset, not the 120-episode Opus confirmatory campaign.

All Qwen LLMs One rational One oracle 00.20.40.60.81.0
ArmMean score [95% CI]Deal rate [95% CI]vs all LLM
All Qwen LLMs0.361 [0.137, 0.615]0.389 [0.167, 0.667]reference
One rational0.325 [0.118, 0.531]0.333 [0.111, 0.556]−0.036 [−0.364, 0.292]
One oracle (spoiled ballot — see erratum)0.810 [0.639, 0.941]0.833 [0.667, 1.000]+0.449 [0.209, 0.692]

Intervals use 10,000 bootstrap samples clustered by the nine parameter sets, keeping both seeds together. Every episode page includes turn-level rational and oracle counterfactual annotations.