Five-seat private frontier campaign (complete)
The main campaign per the “what I want to see” brief: five-seat scorable negotiation, private score sheets, Claude Opus, thinking on, many rollouts, plus its preregistered robustness subsets and a fairness extension. Full methodology and verdicts: research notes 0030/0032/0034/0035 and 0036 in the repo.
- five-seat-opus-ten-arm-complete: the fairness extension — the five arms below plus five more seating agents that are exactly as capable but want the table’s welfare rather than their own, with the decision-surface control that makes the objective swap one-variable. Headline: at one seat out of five, changing only what the computable agent wants is worth
+0.389 utilitarian score and +0.425 deal rate to the whole table (SPOILED VINTAGE — inflated, see the erratum below) — and yet that single fairness seat leaves every welfare aggregate exactly where it found it (Nash welfare, Gini, worst-off share all indistinguishable from the all-LLM table), moving only the geometry of the deals that close. Per-episode pages for both new Opus arms. ⚠ Erratum (2026-08-10): the +0.389 headline is inflated by a spoiled-ballot defect in the two OmniscientBestResponsePolicy arms — that seat silently abstained on its forced-final vote, so one_oracle (the arm the +0.389 is anchored against) is measured on a broken agent. Repaired, one_oracle − all_llm is −0.068 [−0.136, −0.001] rather than −0.412 and all_oracle closes every game (0.875 → 1.000); the accompanying “information amplifies motive” claim is withdrawn (repaired interaction +0.019 [−0.052, +0.091], contains zero). The fairness arms, rational arms, all-LLM and the DP control are clean, and both surviving findings — the single fairness seat’s null on every welfare aggregate, and the omniscient lineup’s worse split — are unaffected. Details on the hub page and in research notes 0039/0045/0043.
- five-seat-opus-five-arm-complete: the main campaign hub — all five arms (all-LLM, one-rational, oracle variants) with per-arm pages, cross-arm analysis, and figures. Private info, datacenter framing, Opus, thinking on. ⚠ Erratum (2026-08-10): the two oracle arms’ closure numbers are the same spoiled ballot — repaired,
all_oracle closes 1.000 rather than 0.875 and its paired score flips −0.082 → +0.034; the erratum banner on the hub carries the full account.
- five-seat-opus-all-llm-vs-one-rational-complete: the paired comparison — all-Opus tables vs the same tables with one computable rational (Bayesian) seat, matched on game and seed.
- five-seat-robustness-opus-abstract-complete: robustness subset 1 — same design with no datacenter framing (abstract issue/option labels), still private-info Opus. ⚠ Erratum (2026-08-10): its
one_oracle arm carries the same spoiled-ballot defect (the omniscient seat silently abstained on its forced-final vote); that arm’s numbers are measured on a defective agent and the subset has not been re-run.
- five-seat-qwen3-8b-robustness-complete: robustness subset 2 — same design on Qwen3-8B (open-weight), private info. ⚠ Erratum (2026-08-10): the
+0.449 one-oracle result carries the spoiled-ballot defect — the omniscient seat silently abstained on its forced-final vote, and this subset has not been re-run, so no corrected figure is available.
Putting the best deal on the table before anyone speaks (August 2026)
- seeded-optimal-opening: the five-arm campaign left open whether the computable lineups’ closure failures are a search failure or a judgment failure, so this experiment deletes the search problem — a neutral non-voting facilitator tables the instance’s ceiling deal (the package the campaign’s normalized score calls 1.0) before any seat speaks, with a mediocre-but-signable placebo as the control. Headline: it fixes nobody’s closure problem. The private Bayesian lineup closes 0.233 of its games with the optimum standing all game, against 0.233 with nothing tabled (paired 0.000 [−0.050, +0.050]). Asked yes-or-no on a signable package, all three lineups behave identically: Opus, Bayesian and omniscient seats all accept in 1.000 of games — for the optimum and for the mediocre placebo alike, down to seats that gain exactly nothing — so the acceptance is signability, not recognition. The two computable families refuse for different reasons when refusing is still possible: Bayesian seats by how thin their own slice is, omniscient seats by continuation value, so the weakest seats accept most readily. And the gate that checked all this found the omniscient seat silently abstaining at the forced final vote in the published campaign — repairing its ballot moves that arm’s deal rate from 0.875 to 1.000 and flips its paired score against all-LLM from −0.082 to +0.034, while leaving the distributional finding intact. Research note 0045; the Opus full-game cell is still in flight.
Two ways to score a perfect fairness number without negotiating (August 2026)
- grpo-v2-lam1-two-attractors: the two degenerate endpoints of one RL run — GRPO on an engine-computed, text-blind log-Nash reward, λ=1.0, fifty steps, evaluated on held-out games at seven checkpoints. It passes through two different degenerate policies, and on the canonical ultimatum holdout both report a below-threshold rate of exactly 0.000 — the guard meant to certify that no seat is pushed under its walk-away. At checkpoint 25 that number is clean because the policy closes 1.000 of its ultimatum episodes while proposing to keep the entire pie: all 15 of 15 episodes land on the same 100/0 split (max share 1.000 against a base of 0.733, where the base model split evenly in 8 of 15), leaving the responder exactly on its threshold — the transcripts have it in the model’s own words, “P1 is exactly my threshold. Accepting P1 closes the deal.” At checkpoint 50 it is clean for the opposite reason: the policy closes nothing at all against a base of 1.000, and the mechanism in the transcripts is not walking away but ceasing to act — in 14 of 15 episodes the proposer emits a no-op instead of tabling any offer. Only the pair {deal-rate floor, share term} separates them: the floor is blind to checkpoint 25 (deal rate a perfect 1.000) and the share term is undefined at checkpoint 50 (no closed deals to take a share of), and no closure-conditional metric sees either. That is the finding behind the program-wide rule that any below-threshold gate carries a share/dispersion term beside its viability floor. Held-out deal rate falls monotonically the whole way (−0.052 → −0.143 → −0.147 → −0.226 → −0.453 → −0.641 → −0.746) while the below-threshold rate stays favourable throughout — fairness bought by refusing to play. Every closure-conditional number at both rungs is stamped VOID by the preregistered viability floor and printed struck-through rather than dropped. Research note 0028; full episode transcripts for every cell.
Earlier runs
- apibehav_mixed_rat0_sonnet5_thinkon: Sonnet negotiation where one out of the five agents is a rational agent and the other four are normal LLMs. Note the following:
- The agents are not told this is a data center construction scenario: they are just given the opaque labels “issue0”, “opt0” and so on
- There are both PRIVATE and FULL runs. In the FULL runs, every agent (including the rational agent) has access to the would-be-private preferences and thresholds of the other agents.
- apibehav_sonnet5_thinkon_cot_datacenter: Sonnet, datacenter scenario, thinking on, no rational agent involved, private info kept private
- five-seat-live-preview-v5: early live preview of the five-seat campaign (Sonnet, no rational agent, private info) — superseded by the complete hubs above
- p2_Qwen3-8B_all_llm_b2: Qwen 8B, many rollouts, no rational agent involved, private info kept private (clean
_b2 re-run, re-rendered with the current viewer)
- grpoeval_lam1_step24_xgame_private: eval rollouts of the fairness-GRPO λ=1 checkpoint on unseen private-info games
De-contamination and the clean re-baseline (July 2026)
The original P2 open-weight campaign was contaminated by a harness bug: on swallowed GPU errors the batched engine silently fabricated placeholder turns that parsed as clean no-ops (26.2% of all turns; up to 100% of single cells). The full story is section B7 of the writeup and research notes 0015/0016; these pages show it and the corrected results.
Writing into the attention mechanism instead of the prompt (August 2026)
A known-useful payload — the exact Nash-bargaining candidate package, worth +0.2741 normalized Nash welfare as ordinary prompt text — delivered instead as per-layer, per-head keys and values at 38 reserved positions. The injected arms and the no-advice control receive a byte-identical prompt, so nothing in these transcripts shows you the intervention; you can only see what it did. Full record: research notes 0029 (rungs R1 and P) and 0046 (rungs U1 and G), plus §7 of the lane writeup.
- qkv_u1_parity_b1: the program’s headline demo — five arms, 1,200 episodes on Qwen3-8B, 0 of 21,606 turns fabricated. For five rungs the channel was reconstruction-perfect but under-acted-on: the model could restate the injected deal on 96–100% of held-out games and still tabled it three times less often than the same deal read as prose, and nothing on the injection side (depth, width, magnitude, a pointer line, the injection site) closed that gap. Retraining the encoder’s objective — supervise the opening-proposal action rather than restatability — does: the channel is worth +0.2232 [+0.1792, +0.2670] against no advice and −0.0386 [−0.0829, +0.0075] against prompt text, i.e. 0.853 [0.706, 1.030] of this panel’s text ceiling against rung R1’s 0.69. Stated at the strength the interval supports, that is parity neither excluded nor established — the registered bar was ≥0.90, the point estimate is below it and the interval contains it. Five paired views, including the one that shows why tabling is not an isolation endpoint (the mis-drawn twin opens on its own wrong package at 29.6% against the treatment arm’s 35.4%, while welfare separates them by +0.1348) and the formatting cost the objective buys (21.7% of episodes malformed against 6.3%, at a site whose R1 encoder costs nothing detectable).
- qkv_g_generality_b1: the generality rung — rung R1’s recipe re-run verbatim (only
--model changed) on Olmo-3-7B-Instruct and Qwen3-4B, 960 episodes each, one panel per model with its own text anchor and no-advice control. Both pass all four preregistered gates, each behind a mechanics gate that proves on the real GPU that re-injecting a span’s own captured K/V leaves the logits bit-identical — including on an architecture with no grouped-query attention and a hybrid sliding/full attention stack. Olmo returns +0.0771 [+0.0505, +0.1022] primary and is statistically indistinguishable from its own text arm (−0.0031 [−0.0343, +0.0273]) — against a weak ceiling of +0.0801, which is the point: across three models the channel/text ratio runs inverse to how well the model follows prose at all. Qwen3-4B returns +0.0830 [+0.0344, +0.1309] at 0.4169 of its own ceiling, and supplies the rung’s oddest finding, which has a view of its own: at 4B prompt text helps while barely tabling anything (15.8% against 62.5% at 8B, a fourfold collapse, for only about a quarter less welfare), so the act that carried the entire text effect at 8B is not what carries it at 4B.
- qkv_r1_kv_vs_text_b1: the first positive — four arms × 24 parameter sets × 10 seeds = 960 episodes on Qwen3-8B, 0 of 17,717 turns fabricated. Per-layer K/V injection is worth +0.1882 [+0.1497, +0.2266] against the token-matched no-advice arm, 69% of the prompt-text ceiling, where the identical information delivered as input-level virtual tokens was a measured zero. It is content and not capacity: the same encoder fed a deliberately mis-drawn deal does significantly worse (+0.0930 [+0.0474, +0.1390]). Three paired views — against no advice, against text, against shuffled content — with the act the whole effect runs through visible at turn 1: opening-offer tabling of the delivered candidate runs 62.1% text / 20.4% K/V / 0% no-advice.
- qkv_p_placement_b1: the placement rung — five arms, 1,200 episodes, asking whether the channel works because the model’s queries already look at the site being written. Writing the payload at real advice text still works (+0.1429 [+0.1008, +0.1855]) but is worse than writing it at a reserved placeholder (−0.0373 [−0.0723, −0.0029]), so site salience is eliminated and the original configuration is locally optimal among everything tested. Four paired views, including the arm’s formatting cost: it ends 30 percentage points [+20.4, +40.4] more episodes malformed than its own control.
Each landing page carries its arms, its per-arm rates and its headline contrasts recomputed from the campaign’s episode table through the campaign’s own instance-cluster bootstrap, checked against the frozen results.json rather than transcribed from it, and links to the other three so the ladder can be walked from any of them.
Paper and data