Say what you want and the table closes; keep guessing and it tops out at 0.41
The published Bayesian negotiator closes 0.233 of these games while the omniscient reference closes 1.000 on the identical bank, so every failure is a package all five seats would have signed, left unsigned. The audit found why: the agent's posterior learns only from a counterpart conceding, these agents almost never concede, and the posterior therefore stays near-uniform — no seat can ever certify that a package would pass. This hub is the confirmation run of the design loop that fixed it. Eleven designs, each a subclass of the frozen code, were scored on a fixed ten-instance subset; the two best were then re-run on the whole 24-instance bank at 120 episodes each, which is what these pages show.
The full-bank confirmation
| arm | deal rate | normalized score | USW | min-seat surplus | IR violations | closing round | IR-for-all proposals round 1 → rounds 3+ | episode pages |
|---|---|---|---|---|---|---|---|---|
all_oracleomniscient ceiling — shown every sheet | 1.000 [1.000, 1.000] | 0.888 [0.868, 0.908] | 193.3 [179.1, 208.6] | 8.40 [6.73, 9.98] | 0.000 [0.000, 0.000] | 1.40 [1.20, 1.63] | 0.932 → 1.000 | reference cell, published in the deadline sweep |
declaringpreference sharing — publish your true sheet, read everyone else's | 1.000 [1.000, 1.000] | 0.857 [0.823, 0.890] | 186.3 [170.5, 203.3] | 7.70 [5.66, 9.83] | 0.000 [0.000, 0.000] | 4.97 [4.95, 5.00] | 0.313 → 1.000 | 10 pages |
conceding_novotesbest design without revelation — concede on a deadline clock | 0.408 [0.283, 0.542] | 0.339 [0.229, 0.453] | 74.8 [49.0, 103.6] | 3.55 [2.21, 5.03] | 0.000 [0.000, 0.000] | 5.00 [5.00, 5.00] | 0.104 → 0.267 | 10 pages |
all_rationalthe published Bayesian agent, unmodified | 0.233 [0.142, 0.333] | 0.195 [0.120, 0.280] | 45.4 [26.6, 66.7] | 3.04 [1.62, 4.68] | 0.000 [0.000, 0.000] | 4.99 [4.97, 5.00] | 0.114 → 0.174 | reference cell, published in the deadline sweep |
120 episodes per arm (24 instances × 5 seeds), four rounds, unanimity, identical
protocol knobs; all four rows scored by the same summarizer
(agent_variants.run_confirmation.summarize), the two reference rows computed at build time from
the deadline sweep's own 4-round cells. Intervals are 95% instance-clustered bootstraps. Min-seat
surplus is the worst-off seat's realized surplus, which is why a package sitting at a seat's exact
walk-away value still counts as a win. IR-for-all proposals is the audit's falsifiable early
indicator — the share of tabled packages acceptable to every seat — and it bends upward exactly
where closure does.
What is not published here. Both confirmation cells are 120 episodes and every number above is measured on all of them; ten episodes per cell have per-episode viewers, rationed against this site's serving ceiling. The complete corpus — 530 episodes and 11,050 turns across the two confirmation cells, the subset bench run and both reference cells — is public in 2026.RA.Agent-Variants-Closure.
Everything behind the page
Full methodology, two corrections (the audit's own predicted fix underdelivered at +0.040
against a predicted 0.5–0.8, and the levers do not compose) and a methods near-miss that nearly produced
a false negative: research note 0059 in the ii_mats repository. Every variant subclasses the
frozen agent; no negotiation solver, policy, scenario or campaign-runner file was edited.