The seat that reads the room

One seat of a five-party unanimity negotiation is given a closed advice loop. Each of its turns, a separate model call reads only the public record — the chat everyone can see and the formal offer ledger — and returns structured claims about what each other party wants, every claim carrying the verbatim sentence it was read off. Those claims are folded into a planner's revealed-preference ledger beside the formal moves, the planner re-ranks candidate packages by the estimated willingness of the least willing party, and the ranked list is appended to that seat's private prompt as explicitly fallible advice. The model keeps the decision, and does its own persuading in public. Research note 0068; the arm is advised_v2__interpreter_advised_llm.

What these pages add that a results table cannot. Every advised turn carries a panel with the loop laid open in the order it ran: the formal-move evidence the planner had per opponent beside the chat-derived claims and their provenance quotes; the ranked candidates it produced; and the move the seat actually played, marked against the ranking. Turns where the seat did not play any ranked candidate are marked distinctly — on the card, on the scrubber chip, and in the audit table above the transcript — because those turns measure the model's own choice rather than the advisor's policy, and roughly half of them are.

What the arm did

Every figure here is read out of the analysis' own summary files at build time, so a re-analysis and a rebuild cannot leave a stale number on this page. The full account, its preregistered gates and its caveats are research note 0068.

Reading the room is worth a resolved increment over speaking alone, and it closes the last of the informed seat's deficit. Against wave 1's arm — the same decision rule with a voice but no ears — this seat's normalized primary is +0.069 [+0.004, +0.155], an interval that excludes zero, and the deals it reaches are more equal (-0.017 [-0.035, -0.002] normalized Gini). Against an ordinary five-Opus table it now closes exactly as often: deal rate +0.000 [-0.042, +0.042]. The seat does not keep the gain — see the override direction below, and note 0068's per-seat capture anchors.

Where the arm lands

armdeal ratenormalized primarynormalized Gini
one_rational — informed, private, and MUTE0.767 [0.683, 0.842]0.686 [0.613, 0.761]0.231 [0.209, 0.255]
A2 — the same rule, given a voice (wave 1)0.900 [0.808, 0.975]0.811 [0.726, 0.884]0.232 [0.210, 0.253]
A5 — one seat reads the room as well (wave 2)0.958 [0.917, 0.992]0.881 [0.836, 0.921]0.214 [0.196, 0.233]
all_llm — five ordinary Opus seats0.958 [0.925, 0.983]0.873 [0.833, 0.910]0.223 [0.200, 0.247]

Levels on the same 24-instance bank, cluster bootstrap over instances, 10000 draws. Lower Gini is more equal.

The paired contrasts

this arm minus...− A2− all_llm
deal rate+0.058 [-0.008, +0.142]+0.000 [-0.042, +0.042]
normalized primary score+0.069 [+0.004, +0.155]+0.008 [-0.034, +0.050]
normalized Nash welfare+0.047 [-0.008, +0.116]+0.007 [-0.031, +0.045]
normalized Gini (lower is more equal)-0.017 [-0.035, -0.002]-0.009 [-0.020, +0.001]
the worst-off party's normalized surplus+0.018 [-0.007, +0.043]+0.011 [-0.011, +0.034]
accepts below a party's own threshold+0.000 [+0.000, +0.000]+0.000 [+0.000, +0.000]

Paired on (instance, seed), cluster bootstrap over the 24 instances. An interval that excludes zero is resolved; one that straddles it is not, however suggestive the point estimate.

Who the gain goes to

this arm − its reference, per seatpaired effect
the four CO-PLAYERS' surplus, unconditional +0.053 [+0.004, +0.120] resolved
the advised seat's own capture -0.072 [-0.267, +0.111]

Read the capture row as a power statement, not a finding. This design cannot resolve the focal-capture endpoint at 120 episodes — wave 1 measured its half-width at 0.227 against a 0.05 threshold — so an interval straddling zero there says the experiment could not tell, not that the effect is absent. What IS resolved is the co-player row: the surplus this better-informed seat unlocks lands on the other four. The override direction below is the mechanism, one turn at a time.

Was the sensor any good?

direction accuracy of the parsed claims 0.886 [0.874, 0.899] against a 0.500 chance baseline, over 8487 scored claims
claims whose quote was findable nowhere in the public text 0.00042 of 9609 rows emitted
share of the planner's evidence that came from chat rather than formal moves 0.801
what the extra parse calls cost $0.48 per episode, 3.69 calls

This is what makes the contrast interpretable in either direction: a null from a sensor that could not read the room would say nothing about whether reading the room helps. The per-turn panels below are the same claims, one turn at a time, each next to the sentence it was read off.

Did the seat take the advice?

An arm whose advice the model discarded measured a prompt, not an intervention, so this is the number that licenses reading any other number the arm produced.

the advice recordcountshare
advised turns452over 120 episodes, 24 instances
turns that played a ranked candidate 243 53.8%
turns that overrode the ranking 209 46.2%
episodes with at least one override101 84.2%
parsed claims that entered the ledger 9604across the traced corpus
Which rung was played, as a share of the 443 turns that were OFFERED a ranked list — a forced final offers no ranking, so this denominator is smaller than the advised-turn count above and the two override shares differ for that reason alone.
played no listed candidate20947.2%
took rank 112428.0%
took rank 24710.6%
took rank 3347.7%
took rank 4296.5%

Read out of advice_trace.json at build time. Compliance is advice_uptake's own verdict — a propose matches on canonical deal identity, an accept matches the offer id a candidate records as already tabled — and neither this page nor the episode viewer re-decides it.

What an override actually is

"Overrode the advice" covers two different behaviours and it is worth separating them before reading the rate above as disagreement. Roughly half of the overrides are the seat accepting a package already on the table that the planner had not ranked — closure, not argument. The other half are the seat tabling a package of its own.

what an overriding turn didturns
accepted a standing offer no candidate named105
proposed a package the planner did not rank88
rejected the offer on the table11
made no formal move5
of the 88 propose-overrides: conceded own surplus / level / moved toward its own optimum 68 / 15 / 5
median change in the seat's own points against the planner's top pick -21.0

The surplus comparison is defined only where the seat proposed a package of its own — an accept names an offer id, not a deal — so its denominator is the propose-overrides, not every override. Read the direction as descriptive: these turns are not a controlled contrast, and the planner's top pick is a capture-leaning candidate by construction, so a seat that closes a deal will usually score below it.

The forced-final branch is exercised here: 9 forced-final turns across 9 episodes, every one of them advised. All 9 issued no parse call, which is the router declining to buy a request it would ignore — the vote is a function of the seat's own sheet alone. The advice on them was 8 accept, 1 reject. Worth stating because a corpus can pass every gate without ever reaching this code: the wave-3 smoke set contains no forced final at all.

The overrides do not run toward the seat's own optimum. Where the comparison is defined — the turns on which the seat proposed its own package — that package is worth less to the seat than the planner's top pick on all but a handful of turns. Whatever these turns are, they are not the seat quietly maximizing against its advisor. The per-turn panels are where to look at what they are instead: the public message the seat published sits directly beneath the advice it did not take.

How a disobeyed turn is identified

Entirely from records that already existed, with no model call and no classifier anywhere in the chain. The planner's ranked advice is recovered from the bytes of the prompt the seat was shown (the advice block is part of the stored view, so it cannot drift from what was actually shown); the move played is the turn's own parsed action; and the comparison between them is advice_uptake's, the module the arm's compliance gate is evaluated on. A turn is marked as an override exactly when that census records no match — the seat proposed a package the planner did not rank, or accepted an offer no candidate names, or did something else entirely.

Two things are deliberately not claimed. The page does not judge whether the seat's public message described its advice faithfully: that is not in the record, and manufacturing it with a classifier would put an unaudited model judgement on a page whose whole purpose is auditing. The advice and the published message are placed together and the reading is left to the reader. Nor does the viewer re-derive compliance; it paints the stored verdict, so the page and the arm's gate cannot disagree.

Episodes you can open

25 of 120 traced episodes are published as full interactive pages. The selection rule is fixed and printed against each row: every episode that did not close, plus the two most- and two least-overridden episodes for each of the five focal seats. The other 95 are not on this site — the corpus is Opus episodes with full reasoning traces and stored prompt views, which is about a hundred megabytes of pages. The audit is complete regardless: analysis/advice_trace.json carries every advised turn of all 120 episodes — the parsed claims with their quotes, the ranked advice, the emitted action and the verdict — so an episode whose page is absent can still be checked.

episodeadvised seat(s)outcometurns overridden claims in the ledgerwhy it is published
00d9a41df9Blakedeal3 / 4116most overridden turns for this focal seat
02041bb0c4Emberdeal0 / 354fewest overridden turns for this focal seat
02bada7c80Blakedeal0 / 356fewest overridden turns for this focal seat
0e65302b5fBlakedeal0 / 596fewest overridden turns for this focal seat
1b417dd7c9Devondeal0 / 354fewest overridden turns for this focal seat
1dbd82a7daEmberdeal0 / 364fewest overridden turns for this focal seat
1ef2fc4adeAverydeal0 / 465fewest overridden turns for this focal seat
273a3ac844Caseydeal3 / 496most overridden turns for this focal seat
287728c7aaCaseydeal0 / 474fewest overridden turns for this focal seat
2e51134450Devondeal3 / 498most overridden turns for this focal seat
3008eca3a5Caseyno deal3 / 480did not close
33d2ddda15Emberdeal4 / 4113most overridden turns for this focal seat
49580093f9Caseydeal4 / 494most overridden turns for this focal seat
5efa29acefAverydeal0 / 483fewest overridden turns for this focal seat
65223d5cdaEmberno deal3 / 5128did not close
6fbc8b84b6Caseyno deal1 / 5101did not close
78b9e83092Blakedeal4 / 4116most overridden turns for this focal seat
98c77d55e8Caseyno deal3 / 5106did not close
9b0fb718c2Emberdeal4 / 481most overridden turns for this focal seat
a15fefb0afCaseydeal0 / 345fewest overridden turns for this focal seat
a21ac7267dDevondeal0 / 348fewest overridden turns for this focal seat
cad30bc706Blakeno deal2 / 575did not close
d0ec680977Averydeal4 / 483most overridden turns for this focal seat
d6107aede5Devondeal4 / 482most overridden turns for this focal seat
d9ea2847dcAverydeal4 / 488most overridden turns for this focal seat

The episode index for the published set: sortable table of all 25 pages.

Artifacts

All 39 analysis file(s) are under analysis/, including the per-claim CSVs and the fairness-basket figures. Run directory /nlp/scr/siddharth/ii_mats/rational_agents/advised_v2__interpreter_advised_llm; trace built 2026-08-20T20:50:12.732171+00:00, schema five-seat-interpreter-advice-trace-v1. Every advised turn joined its parse sidecar (0 unjoined).