When everyone reads the room

Every one of the five seats of a five-party unanimity negotiation is given a closed advice loop. Each of its turns, a separate model call reads only the public record — the chat everyone can see and the formal offer ledger — and returns structured claims about what each other party wants, every claim carrying the verbatim sentence it was read off. Those claims are folded into a planner's revealed-preference ledger beside the formal moves, the planner re-ranks candidate packages by the estimated willingness of the least willing party, and the ranked list is appended to that seat's private prompt as explicitly fallible advice. The model keeps the decision, and does its own persuading in public. Each seat has its own parser, its own planner keyed to its own score sheet, and its own claim history: there is no shared ledger, so no seat's evidence can reach another's plan. That isolation is what keeps this from being an omniscient arm wearing an advised arm's name. Research note 0075; the arm is advised_v3__all_interpreter_advised_llm.

What these pages add that a results table cannot. Every advised turn carries a panel with the loop laid open in the order it ran: the formal-move evidence the planner had per opponent beside the chat-derived claims and their provenance quotes; the ranked candidates it produced; and the move the seat actually played, marked against the ranking. Turns where the seat did not play any ranked candidate are marked distinctly — on the card, on the scrubber chip, and in the audit table above the transcript — because those turns measure the model's own choice rather than the advisor's policy, and roughly half of them are. On this arm every seat is advised, so each round is five such decisions at once against one table state — the round-by-round view below lays them out five-up.

What the arm did

Every figure here is read out of the analysis' own summary files at build time, so a re-analysis and a rebuild cannot leave a stale number on this page. The full account, its preregistered gates and its caveats are research note 0075.

Every seat concedes, and the table ends up less equal. When one seat reads the room it spends its own surplus to buy closure and the other four collect it. When ALL FIVE do, each is still a donor — every seat individually — and yet against the one-donor arm the agreements are less equal, not more: normalized Gini +0.015 [+0.005, +0.025] (higher is less equal) and the worst-off party's surplus -0.030 [-0.053, -0.008], both intervals excluding zero, while the primary score +0.018 [-0.010, +0.050] does not resolve. Those are the facts and they do not depend on any account of the mechanism — and the mechanism is an open question, not a finding. The obvious story, that simultaneous concessions cancel where a lone concession transfers, was written down as a prediction with a falsifier and then tested: both prespecified correlations came back null (conceder count against realized Gini, and concentration of conceding against it). So the round-by-round view below is where to LOOK, not an explanation — it is the substrate a future arm would settle this against, and settling it needs a design that ASSIGNS how many seats are advised rather than observing it.
Which contrast is the preregistered primary, and which one carries the finding — they are not the same contrast. The registered primary is against the uninformed baseline, and it comes back null on everything: against five ordinary Opus seats this arm reads +0.026 [-0.010, +0.066] on primary, +0.005 [-0.007, +0.017] on Gini and -0.015 [-0.044, +0.013] on the worst-off party, every interval straddling zero. So the honest reading is null against the uninformed baseline, adverse against the one-donor arm — and the adverse A6 − A5 contrast above, which is where the resolved intervals are, was not the registered primary. It is shown first because it is the comparison this arm exists to draw, not because it is the one that was registered.

Where the arm lands

armdeal ratenormalized primarynormalized Gini
one_rational — informed, private, and MUTE0.767 [0.683, 0.842]0.686 [0.614, 0.758]0.231 [0.209, 0.255]
A2 — the same rule, given a voice (wave 1)0.900 [0.808, 0.975]0.811 [0.726, 0.884]0.232 [0.209, 0.253]
A5 — one seat reads the room as well (wave 2)0.958 [0.917, 0.992]0.881 [0.836, 0.921]0.214 [0.196, 0.233]
A6 — every seat reads the room (wave 3)0.983 [0.950, 1.000]0.899 [0.856, 0.935]0.229 [0.207, 0.250]
all_llm — five ordinary Opus seats0.958 [0.925, 0.992]0.873 [0.833, 0.910]0.223 [0.200, 0.247]

Levels on the same 24-instance bank, cluster bootstrap over instances, 10000 draws. Lower Gini is more equal.

The paired contrasts

this arm minus...− A5− all_llm− all_rational
deal rate+0.025 [+0.000, +0.050]+0.025 [-0.008, +0.058]+0.750 [+0.633, +0.850]
normalized primary score+0.018 [-0.010, +0.050]+0.026 [-0.010, +0.066]+0.710 [+0.602, +0.810]
normalized Nash welfare-0.018 [-0.041, +0.007]-0.010 [-0.056, +0.027]+0.407 [+0.320, +0.491]
normalized Gini (lower is more equal)+0.015 [+0.005, +0.025]+0.005 [-0.007, +0.017]-0.025 [-0.072, +0.019]
the worst-off party's normalized surplus-0.030 [-0.053, -0.008]-0.015 [-0.044, +0.013]+0.114 [+0.033, +0.199]
accepts below a party's own threshold+0.000 [+0.000, +0.000]+0.000 [+0.000, +0.000]+0.000 [+0.000, +0.000]

Paired on (instance, seed), cluster bootstrap over the 24 instances. An interval that excludes zero is resolved; one that straddles it is not, however suggestive the point estimate.

Is every seat a donor?

seatuptakeconcede / level / grabconcession rate median own-surplus change
Avery0.53469 / 14 / 10.821 [0.743, 0.889]-27.0
Blake0.55742 / 21 / 30.636 [0.448, 0.774]-19.0
Casey0.56369 / 8 / 10.885 [0.815, 0.947]-23.5
Devon0.57057 / 18 / 50.713 [0.597, 0.810]-24.0
Ember0.59947 / 15 / 100.653 [0.517, 0.768]-17.0

Counted over the overrides where the comparison is defined — a seat that proposed a package of its own, since an accept names an offer id and not a deal. There is no focal-capture row on this arm and its absence is by design: within-table capture is only defined against untreated co-players, and this table has none.

Was the sensor any good?

direction accuracy of the parsed claims 0.873 [0.861, 0.885] against a 0.500 chance baseline, over 44937 scored claims
claims whose quote was findable nowhere in the public text 0.00056 of 51984 rows emitted
share of the planner's evidence that came from chat rather than formal moves 0.814
what the extra parse calls cost $2.77 per episode, 18.56 calls

This is what makes the contrast interpretable in either direction: a null from a sensor that could not read the room would say nothing about whether reading the room helps. The per-turn panels below are the same claims, one turn at a time, each next to the sentence it was read off.

Did the seat take the advice?

An arm whose advice the model discarded measured a prompt, not an intervention, so this is the number that licenses reading any other number the arm produced.

the advice recordcountshare
advised turns2263over 120 episodes, 24 instances
turns that played a ranked candidate 1277 56.4%
turns that overrode the ranking 986 43.6%
episodes with at least one override120 100.0%
parsed claims that entered the ledger 51950across the traced corpus
Which rung was played, as a share of the 2227 turns that were OFFERED a ranked list — a forced final offers no ranking, so this denominator is smaller than the advised-turn count above and the two override shares differ for that reason alone.
played no listed candidate98544.2%
took rank 161327.5%
took rank 233114.9%
took rank 31567.0%
took rank 41426.4%

Read out of advice_trace.json at build time. Compliance is advice_uptake's own verdict — a propose matches on canonical deal identity, an accept matches the offer id a candidate records as already tabled — and neither this page nor the episode viewer re-decides it.

What an override actually is

"Overrode the advice" covers two different behaviours and it is worth separating them before reading the rate above as disagreement. Roughly half of the overrides are the seat accepting a package already on the table that the planner had not ranked — closure, not argument. The other half are the seat tabling a package of its own.

what an overriding turn didturns
accepted a standing offer no candidate named511
proposed a package the planner did not rank380
rejected the offer on the table69
made no formal move26
of the 380 propose-overrides: conceded own surplus / level / moved toward its own optimum 284 / 76 / 20
median change in the seat's own points against the planner's top pick -22.0

The surplus comparison is defined only where the seat proposed a package of its own — an accept names an offer id, not a deal — so its denominator is the propose-overrides, not every override. Read the direction as descriptive: these turns are not a controlled contrast, and the planner's top pick is a capture-leaning candidate by construction, so a seat that closes a deal will usually score below it.

The forced-final branch is exercised here: 36 forced-final turns across 9 episodes, every one of them advised. All 36 issued no parse call, which is the router declining to buy a request it would ignore — the vote is a function of the seat's own sheet alone. The advice on them was 35 accept, 1 reject. Worth stating because a corpus can pass every gate without ever reaching this code: the wave-3 smoke set contains no forced final at all.

The overrides do not run toward the seat's own optimum. Where the comparison is defined — the turns on which the seat proposed its own package — that package is worth less to the seat than the planner's top pick on all but a handful of turns. Whatever these turns are, they are not the seat quietly maximizing against its advisor. The per-turn panels are where to look at what they are instead: the public message the seat published sits directly beneath the advice it did not take.

How a disobeyed turn is identified

Entirely from records that already existed, with no model call and no classifier anywhere in the chain. The planner's ranked advice is recovered from the bytes of the prompt the seat was shown (the advice block is part of the stored view, so it cannot drift from what was actually shown); the move played is the turn's own parsed action; and the comparison between them is advice_uptake's, the module the arm's compliance gate is evaluated on. A turn is marked as an override exactly when that census records no match — the seat proposed a package the planner did not rank, or accepted an offer no candidate names, or did something else entirely.

Two things are deliberately not claimed. The page does not judge whether the seat's public message described its advice faithfully: that is not in the record, and manufacturing it with a classifier would put an unaudited model judgement on a page whose whole purpose is auditing. The advice and the published message are placed together and the reading is left to the reader. Nor does the viewer re-derive compliance; it paints the stored verdict, so the page and the arm's gate cannot disagree.

Episodes you can open

22 of 120 traced episodes are published as full interactive pages. The selection rule is fixed and printed against each row: every episode that did not close, plus the two most- and two least-overridden episodes for each of the five focal seats. The other 98 are not on this site — the corpus is Opus episodes with full reasoning traces and stored prompt views, which is about a hundred megabytes of pages. The audit is complete regardless: analysis/advice_trace.json carries every advised turn of all 120 episodes — the parsed claims with their quotes, the ranked advice, the emitted action and the verdict — so an episode whose page is absent can still be checked.

episodeadvised seat(s)outcometurns overridden claims in the ledgerwhy it is published
008fea746call 5deal16 / 20517highest share of turns overriding the advice
01c409ccb8all 5deal4 / 19520lowest share of turns overriding the advice
03c418fa01all 5deal2 / 15337lowest share of turns overriding the advice
0d6fb4f14aall 5deal14 / 20545highest share of turns overriding the advice
1d21110bd2all 5no deal9 / 25529did not close
220940dc51all 5deal15 / 20441highest share of turns overriding the advice
275b93b934all 5deal3 / 18398lowest share of turns overriding the advice
3a8d88dc97all 5deal13 / 17390highest share of turns overriding the advice
3f267f7b99all 5deal3 / 15328lowest share of turns overriding the advice
48a355695call 5deal4 / 20491lowest share of turns overriding the advice
56ebc51223all 5deal16 / 20502highest share of turns overriding the advice
6424f1b8cdall 5no deal7 / 25545did not close
7181763c40all 5deal14 / 20484highest share of turns overriding the advice
832ab38222all 5deal11 / 15330highest share of turns overriding the advice
94cb24d769all 5deal2 / 14224lowest share of turns overriding the advice
9f87bf1f94all 5deal3 / 20482lowest share of turns overriding the advice
ac411850f0all 5deal14 / 20443highest share of turns overriding the advice
b25b5cb4d2all 5deal13 / 19444highest share of turns overriding the advice
bff73cceb9all 5deal10 / 15408highest share of turns overriding the advice
d0f7b6ead0all 5deal3 / 15303lowest share of turns overriding the advice
e458fcc48dall 5deal4 / 20483lowest share of turns overriding the advice
e50748a08eall 5deal4 / 20528lowest share of turns overriding the advice

The episode index for the published set: sortable table of all 22 pages.

Artifacts

All 41 analysis file(s) are under analysis/, including the per-claim CSVs and the fairness-basket figures. Run directory /nlp/scr/siddharth/ii_mats/rational_agents/advised_v3__all_interpreter_advised_llm; trace built 2026-08-25T07:59:32.553921+00:00, schema five-seat-interpreter-advice-trace-v1. Every advised turn joined its parse sidecar (0 unjoined).