Was it the rational agent, or was it the veto?
The frozen five-seat campaign found that replacing one of five Claude Opus 5 seats with a
computable rational agent costs the table -0.186 [-0.266, -0.105] of
normalized score, almost entirely through closure. Every episode of that campaign ran under
unanimity — the rule that gives that single seat a veto. These two
arms re-run both lineups under a 4-of-5 majority rule and change nothing else: same 24
parameter sets, same five seeds, same model, thinking, framing, deadline and rotating proposer. Every
episode pairs with a frozen episode on (instance id, seed), so the estimand is a
difference in differences.
Arms
| Arm | Normalized score [95% CI] | Deal rate [95% CI] | Below-threshold outcome |
|---|---|---|---|
all_llmfive Opus seats, unanimity (5 of 5 must accept) | 0.873 [0.833, 0.912] | 0.958 [0.925, 0.992] | 0.000 [0.000, 0.000] |
one_rational4 Opus + 1 private Bayesian-rational seat, unanimity | 0.686 [0.613, 0.761] | 0.767 [0.683, 0.842] | 0.000 [0.000, 0.000] |
all_llm_quorum_4 this armfive Opus seats, MAJORITY (4 of 5) | 0.924 [0.902, 0.946] | 1.000 [1.000, 1.000] | 0.242 [0.133, 0.358] |
one_rational_quorum_4 this arm4 Opus + 1 private Bayesian-rational seat, MAJORITY (4 of 5) | 0.868 [0.824, 0.909] | 0.983 [0.958, 1.000] | 0.375 [0.258, 0.500] |
A no-deal scores zero. Intervals are 95% cluster bootstraps over the 24 parameter sets, each set's five seeds resampled together.
The difference in differences
The treated seat's cost, measured under each rule, and the difference between them. A DiD that cancels the unanimity column means the cost was the power to block; a DiD indistinguishable from zero means it was not.
| Metric | one_rational − all_llmunder unanimity | the same contrast under 4-of-5 | difference in differences |
|---|---|---|---|
| normalized score (primary) | -0.186 [-0.266, -0.105] | -0.056 [-0.091, -0.025] | +0.130 [+0.053, +0.208] |
| deal rate | -0.192 [-0.275, -0.108] | -0.017 [-0.042, +0.000] | +0.175 [+0.092, +0.258] |
| below-threshold outcome rate | +0.000 [+0.000, +0.000] | +0.133 [+0.025, +0.242] | +0.133 [+0.025, +0.242] |
| normalized Nash welfare | -0.124 [-0.187, -0.065] | -0.089 [-0.139, -0.038] | +0.035 [-0.052, +0.124] |
| worst-off z (among deals) | -0.011 [-0.043, +0.021] | -0.590 [-1.626, -0.048] | -0.348 [-1.102, +0.042] |
| normalized Gini (among deals) | +0.005 [-0.015, +0.026] | +0.076 [+0.037, +0.122] | +0.032 [-0.013, +0.073] |
| max share (among deals) | +0.000 [-0.009, +0.010] | +0.038 [+0.017, +0.063] | +0.018 [-0.008, +0.041] |
What the rule alone did to each lineup
The protocol's own main effect, which separates “the treated seat stopped mattering” from “everything closes more easily once four names suffice”.
| Lineup, majority minus unanimity | Normalized score | Deal rate | Below-threshold outcome |
|---|---|---|---|
all_llm | +0.052 [+0.005, +0.099] | +0.042 [+0.008, +0.075] | +0.242 [+0.133, +0.367] |
one_rational | +0.182 [+0.119, +0.248] | +0.217 [+0.150, +0.292] | +0.375 [+0.258, +0.492] |
The outvoted seat
A quorum close is a deal that passed without every seat's support; the outvoted seat is the seat absent from the closing package's support. Because accepts accumulate across offers in this protocol, silence is not refusal — so the last column separates a seat that formally rejected the package that passed (a deal closed over an objection) from one that simply never acted on it.
| Arm | Deals | Closed over a seat | Outvoted seat below threshold | Outvoted seat's z | Outvoted seat dissented |
|---|---|---|---|---|---|
all_llm | 115 | 0 | — | — | — |
one_rational | 92 | 0 | — | — | — |
all_llm_quorum_4 | 120 | 118 | 0.246 [0.134, 0.370] | 0.164 [0.108, 0.221] | 0.059 [0.017, 0.109] |
one_rational_quorum_4 | 118 | 112 | 0.402 [0.276, 0.527] | 0.056 [-0.010, 0.122] | 0.054 [0.018, 0.091] |
Did the rational seat change its behaviour, or only its leverage?
The section below reconciles this arm with the sibling moves-only control. What separates their two explanations is whether the seat behaves differently when only the price of withholding agreement is removed: an unchanged reject share while the table closes over it means the veto explained the price and something upstream explained the behaviour.
Both columns are shown because the veto is exercisable passively: under unanimity a seat blocks just as effectively by never accepting as by rejecting. On the frozen unanimity arm this seat rejects almost nothing and simply withholds agreement, so a reject-only reading would call it cooperative.
| Arm | Treated seat's reject share | Accept share | Outvoted | Left below threshold |
|---|---|---|---|---|
one_rational | 0.027 [0.015, 0.040] | 0.351 [0.312, 0.391] | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] |
one_rational_quorum_4 | 0.000 [0.000, 0.000] | 0.313 [0.267, 0.358] | 0.542 [0.470, 0.615] | 0.203 [0.134, 0.271] |
How this fits the tools-matched moves-only control
The rational seat differs from an LLM seat in at least two ways at once: it cannot use the message channel, and under unanimity its withholding of agreement is decisive. These sit at different points on one causal chain — no channel, so it never converges with the table's forming consensus, so it withholds agreement late; unanimity, so that lone withholding voids the episode. Channel access is the upstream candidate explanation of why it fails to agree; veto power is the downstream explanation of why that failure costs the full −0.186. Both interventions can therefore each act on the gap without contradicting one another, and neither lane's result subsumes the other's.
The two halves at their measured strengths, deliberately not averaged:
- Upstream, measured by the moves-only control: closing the cheap-talk channel for all five seats reproduces the gap in full — a mute all-Opus table and the one-rational table are statistically indistinguishable, −0.023 [−0.118, +0.064] on normalized score.
- Downstream, measured here: removing the veto erases 70% of the gap (+0.130 [+0.053, +0.208]) and leaves a small but significant residual of -0.056 [-0.091, -0.025]. Partial, and reported as partial.
Two cautions belong with this. Muting all five seats is not the same intervention as the rational seat's asymmetric muteness, so that contrast bounds the value of the channel rather than isolating the asymmetry — which is why channel access stays a candidate upstream explanation and is not promoted to explaining the seat. And the joint cell that would disentangle them, one_rational × quorum_4 × moves_only, has not been run and is named as future work in both notes.
Everything behind the page
all_llm_quorum_4Per-episode transcripts: one_rational_quorum_4
The frozen five-arm campaign
The tools-matched moves-only control
Dataset on Hugging Face
Full preregistration, endpoints and adjudication: research note 0049 (executing note 0047
arm C) in the ii_mats repository. A sibling lane asks the same question about a different piece
of the protocol — whether the published cost is really about which seats were allowed to talk —
and the two readings are reconciled in the notes.