## Five-seat private frontier campaign (complete)

The main campaign per the "what I want to see" brief: **five-seat scorable negotiation, private score sheets, Claude Opus, thinking on, many rollouts**, plus its preregistered robustness subsets and a fairness extension. Full methodology and verdicts: research notes 0030/0032/0034/0035 and 0036 in the repo.

- [five-seat-opus-ten-arm-complete](./five-seat-opus-ten-arm-complete/index.html): the **fairness extension** — the five arms below plus five more seating agents that are exactly as capable but want the *table's* welfare rather than their own, with the decision-surface control that makes the objective swap one-variable. Headline: at one seat out of five, changing only what the computable agent wants is worth ~~**+0.389 utilitarian score and +0.425 deal rate to the whole table**~~ (SPOILED VINTAGE — inflated, see the erratum below) — and yet that single fairness seat leaves **every welfare aggregate exactly where it found it** (Nash welfare, Gini, worst-off share all indistinguishable from the all-LLM table), moving only the geometry of the deals that close. Per-episode pages for both new Opus arms. **⚠ Erratum (2026-08-10): the +0.389 headline is inflated by a spoiled-ballot defect in the two `OmniscientBestResponsePolicy` arms** — that seat silently abstained on its forced-final vote, so `one_oracle` (the arm the +0.389 is anchored against) is measured on a broken agent. Repaired, `one_oracle` − `all_llm` is −0.068 [−0.136, −0.001] rather than −0.412 and `all_oracle` closes every game (0.875 → 1.000); the accompanying "information amplifies motive" claim is **withdrawn** (repaired interaction +0.019 [−0.052, +0.091], contains zero). The fairness arms, rational arms, all-LLM and the DP control are clean, and both surviving findings — the single fairness seat's null on every welfare aggregate, and the omniscient lineup's worse split — are unaffected. Details on the hub page and in research notes 0039/0045/0043.
- [five-seat-opus-five-arm-complete](./five-seat-opus-five-arm-complete/index.html): the **main campaign hub** — all five arms (all-LLM, one-rational, oracle variants) with per-arm pages, cross-arm analysis, and figures. Private info, datacenter framing, Opus, thinking on. **⚠ Erratum (2026-08-10):** the two oracle arms' closure numbers are the same spoiled ballot — repaired, `all_oracle` closes 1.000 rather than 0.875 and its paired score flips −0.082 → +0.034; the erratum banner on the hub carries the full account.
- [five-seat-opus-all-llm-vs-one-rational-complete](./five-seat-opus-all-llm-vs-one-rational-complete/index.html): the paired comparison — all-Opus tables vs the same tables with one computable rational (Bayesian) seat, matched on game and seed.
- [five-seat-robustness-opus-abstract-complete](./five-seat-robustness-opus-abstract-complete/index.html): robustness subset 1 — same design with **no datacenter framing** (abstract issue/option labels), still private-info Opus. **⚠ Erratum (2026-08-10):** its `one_oracle` arm carries the same spoiled-ballot defect (the omniscient seat silently abstained on its forced-final vote); that arm's numbers are measured on a defective agent and the subset has not been re-run.
- [five-seat-qwen3-8b-robustness-complete](./five-seat-qwen3-8b-robustness-complete/index.html): robustness subset 2 — same design on **Qwen3-8B** (open-weight), private info. **⚠ Erratum (2026-08-10):** the `+0.449` one-oracle result carries the spoiled-ballot defect — the omniscient seat silently abstained on its forced-final vote, and this subset has not been re-run, so no corrected figure is available.

<!-- [implement: rational_agents orig (results review)--optimal on the table] 2026-08-10 — the seeded-optimal-opening hub. -->

## Putting the best deal on the table before anyone speaks (August 2026)

- [seeded-optimal-opening](./seeded-optimal-opening/index.html): the five-arm campaign left open whether the computable lineups' closure failures are a **search** failure or a **judgment** failure, so this experiment deletes the search problem — a neutral non-voting facilitator tables the instance's **ceiling deal** (the package the campaign's normalized score calls 1.0) before any seat speaks, with a mediocre-but-signable **placebo** as the control. **Headline: it fixes nobody's closure problem.** The private Bayesian lineup closes **0.233** of its games with the optimum standing all game, against **0.233** with nothing tabled (paired 0.000 [−0.050, +0.050]). Asked yes-or-no on a signable package, **all three lineups behave identically**: Opus, Bayesian and omniscient seats all accept in **1.000** of games — for the optimum and for the mediocre placebo alike, down to seats that gain exactly nothing — so the acceptance is signability, not recognition. The two computable families refuse for different reasons when refusing is still possible: Bayesian seats by how thin their own slice is, omniscient seats by continuation value, so the weakest seats accept most readily. **And the gate that checked all this found the omniscient seat silently abstaining at the forced final vote in the published campaign** — repairing its ballot moves that arm's deal rate from 0.875 to **1.000** and flips its paired score against all-LLM from −0.082 to **+0.034**, while leaving the distributional finding intact. Research note 0045; the Opus full-game cell is still in flight.

<!-- [fix: rational_agents orig (results review)] 2026-08-10 — the fairness-GRPO-v2 λ=1.0 two-attractors hub (checkpoints 25 and 50). Built by experiments/rational_agents/build_grpo_v2_checkpoint_hub.py. -->

## Two ways to score a perfect fairness number without negotiating (August 2026)

- [grpo-v2-lam1-two-attractors](./grpo-v2-lam1-two-attractors/index.html): the two degenerate endpoints of one RL run — GRPO on an engine-computed, text-blind log-Nash reward, λ=1.0, fifty steps, evaluated on held-out games at seven checkpoints. **It passes through two different degenerate policies, and on the canonical ultimatum holdout both report a below-threshold rate of exactly 0.000** — the guard meant to certify that no seat is pushed under its walk-away. At **[checkpoint 25](./grpo-v2-lam1-two-attractors/ckpt25/index.html)** that number is clean because the policy closes **1.000** of its ultimatum episodes while proposing to keep the entire pie: all 15 of 15 episodes land on the same **100/0** split (max share **1.000** against a base of 0.733, where the base model split evenly in 8 of 15), leaving the responder *exactly* on its threshold — the transcripts have it in the model's own words, "P1 is exactly my threshold. Accepting P1 closes the deal." At **[checkpoint 50](./grpo-v2-lam1-two-attractors/ckpt50/index.html)** it is clean for the opposite reason: the policy closes **nothing at all** against a base of 1.000, and the mechanism in the transcripts is not walking away but **ceasing to act** — in 14 of 15 episodes the proposer emits a no-op instead of tabling any offer. **Only the pair {deal-rate floor, share term} separates them**: the floor is blind to checkpoint 25 (deal rate a perfect 1.000) and the share term is undefined at checkpoint 50 (no closed deals to take a share of), and no closure-conditional metric sees either. That is the finding behind the program-wide rule that any below-threshold gate carries a share/dispersion term beside its viability floor. Held-out deal rate falls monotonically the whole way (−0.052 → −0.143 → −0.147 → −0.226 → −0.453 → −0.641 → **−0.746**) while the below-threshold rate stays *favourable* throughout — fairness bought by refusing to play. Every closure-conditional number at both rungs is stamped VOID by the preregistered viability floor and printed struck-through rather than dropped. Research note 0028; full episode transcripts for every cell.

## Earlier runs

- [apibehav_mixed_rat0_sonnet5_thinkon](./apibehav_mixed_rat0_sonnet5_thinkon/index.html): Sonnet negotiation where one out of the five agents is a rational agent and the other four are normal LLMs. Note the following:
	- The agents are not told this is a data center construction scenario: they are just given the opaque labels "issue0", "opt0" and so on
	- There are both PRIVATE and FULL runs. In the FULL runs, every agent (including the rational agent) has access to the would-be-private preferences and thresholds of the other agents.
- [apibehav_sonnet5_thinkon_cot_datacenter](./apibehav_sonnet5_thinkon_cot_datacenter/index.html): Sonnet, datacenter scenario, thinking on, no rational agent involved, private info kept private
- [five-seat-live-preview-v5](./five-seat-live-preview-v5/index.html): early live preview of the five-seat campaign (Sonnet, no rational agent, private info) — superseded by the complete hubs above
- [p2_Qwen3-8B_all_llm_b2](./p2_Qwen3-8B_all_llm_b2/index.html): Qwen 8B, many rollouts, no rational agent involved, private info kept private (clean `_b2` re-run, re-rendered with the current viewer)
- [grpoeval_lam1_step24_xgame_private](./grpoeval_lam1_step24_xgame_private/index.html): eval rollouts of the fairness-GRPO λ=1 checkpoint on unseen private-info games

## De-contamination and the clean re-baseline (July 2026)

The original P2 open-weight campaign was contaminated by a harness bug: on swallowed GPU errors the batched engine silently fabricated placeholder turns that parsed as clean no-ops (26.2% of all turns; up to 100% of single cells). The full story is section B7 of the writeup and research notes 0015/0016; these pages show it and the corrected results.

- [decontamination_Qwen3-8B_all_llm](./decontamination_Qwen3-8B_all_llm/index.html): **the visual proof** — contaminated episodes (red "NOT GENERATED" badges on every fabricated turn) paired against the clean re-run of the same game/seed. Lead pair: 20 of 25 turns never generated, no-deal → deal, welfare 0 → 348.
- Clean seat-swap comparisons (rational agent in the exact same seat, both sides 0.0% fabricated): [seatswap_reverse_rseat1_b2](./seatswap_reverse_rseat1_b2/index.html), [seatswap_reverse_Qwen3-4B_rseat4_b2](./seatswap_reverse_Qwen3-4B_rseat4_b2/index.html), [seatswap_reverse_Qwen3-8B_rseat0_b2](./seatswap_reverse_Qwen3-8B_rseat0_b2/index.html), [seatswap_p2_mixed_b2](./seatswap_p2_mixed_b2/index.html). Post-cleanup verdict: the rational seat's *capture* advantage survives on every slot; the "repairs deal-closing" claim did not (it was mostly the artifact suppressing the baseline).
- [apibehav_sonnet5_thinkoff](./apibehav_sonnet5_thinkoff/index.html) and [apibehav_haiku45_thinkoff](./apibehav_haiku45_thinkoff/index.html): the frontier-API headline cells (thinking off, matched to the open-weight protocol) behind the "Claudes close more deals but distribute them worse" result.
- [ultimatum_giveaway](./ultimatum_giveaway/index.html): the ultimatum-game preset — a one-flag situation swap on the same harness.

<!-- [implement: rational_agents new transformer feature] 2026-08-09 — Agent V, session id 617b6f67-2120-4239-a51c-1c0b0f3a762f. The Q/K/V attention-interface paired viewers. Built by experiments/rational_agents/tom/qkv/build_qkv_paired_viewer.py. -->

## Writing into the attention mechanism instead of the prompt (August 2026)

A known-useful payload — the exact Nash-bargaining candidate package, worth **+0.2741 normalized Nash welfare** as ordinary prompt text — delivered instead as per-layer, per-head keys and values at 38 reserved positions. The injected arms and the no-advice control receive a **byte-identical prompt**, so nothing in these transcripts shows you the intervention; you can only see what it did. Full record: research notes 0029 (rungs R1 and P) and 0046 (rungs U1 and G), plus §7 of the [lane writeup](./writeup_tom_lane.pdf).

<!-- [implement: rational_agents new transformer feature] 2026-08-10 — Agent V2, same session: the U1 parity viewer and the two-model generality viewer. -->

- [qkv_u1_parity_b1](./qkv_u1_parity_b1/index.html): the **program's headline demo** — five arms, 1,200 episodes on Qwen3-8B, 0 of 21,606 turns fabricated. For five rungs the channel was *reconstruction-perfect but under-acted-on*: the model could restate the injected deal on 96–100% of held-out games and still tabled it three times less often than the same deal read as prose, and nothing on the injection side (depth, width, magnitude, a pointer line, the injection site) closed that gap. Retraining the **encoder's objective** — supervise the opening-proposal action rather than restatability — does: the channel is worth **+0.2232 [+0.1792, +0.2670]** against no advice and **−0.0386 [−0.0829, +0.0075]** against prompt text, i.e. **0.853 [0.706, 1.030]** of this panel's text ceiling against rung R1's 0.69. Stated at the strength the interval supports, that is *parity neither excluded nor established* — the registered bar was ≥0.90, the point estimate is below it and the interval contains it. Five paired views, including the one that shows why tabling is not an isolation endpoint (the mis-drawn twin opens on its own **wrong** package at 29.6% against the treatment arm's 35.4%, while welfare separates them by +0.1348) and the formatting cost the objective buys (21.7% of episodes malformed against 6.3%, at a site whose R1 encoder costs nothing detectable).
- [qkv_g_generality_b1](./qkv_g_generality_b1/index.html): the **generality rung** — rung R1's recipe re-run verbatim (only `--model` changed) on **Olmo-3-7B-Instruct** and **Qwen3-4B**, 960 episodes each, one panel per model with its own text anchor and no-advice control. Both pass all four preregistered gates, each behind a mechanics gate that proves on the real GPU that re-injecting a span's own captured K/V leaves the logits **bit-identical** — including on an architecture with no grouped-query attention and a hybrid sliding/full attention stack. Olmo returns **+0.0771 [+0.0505, +0.1022]** primary and is *statistically indistinguishable from its own text arm* (**−0.0031 [−0.0343, +0.0273]**) — against a weak ceiling of +0.0801, which is the point: across three models the channel/text ratio runs inverse to how well the model follows prose at all. Qwen3-4B returns **+0.0830 [+0.0344, +0.1309]** at 0.4169 of its own ceiling, and supplies the rung's oddest finding, which has a view of its own: at 4B **prompt text helps while barely tabling anything** (15.8% against 62.5% at 8B, a fourfold collapse, for only about a quarter less welfare), so the act that carried the entire text effect at 8B is not what carries it at 4B.
- [qkv_r1_kv_vs_text_b1](./qkv_r1_kv_vs_text_b1/index.html): the **first positive** — four arms × 24 parameter sets × 10 seeds = 960 episodes on Qwen3-8B, 0 of 17,717 turns fabricated. Per-layer K/V injection is worth **+0.1882 [+0.1497, +0.2266]** against the token-matched no-advice arm, **69% of the prompt-text ceiling**, where the identical information delivered as input-level virtual tokens was a measured zero. It is content and not capacity: the same encoder fed a deliberately mis-drawn deal does significantly worse (**+0.0930 [+0.0474, +0.1390]**). Three paired views — against no advice, against text, against shuffled content — with the act the whole effect runs through visible at turn 1: opening-offer tabling of the delivered candidate runs **62.1% text / 20.4% K/V / 0% no-advice**.
- [qkv_p_placement_b1](./qkv_p_placement_b1/index.html): the **placement rung** — five arms, 1,200 episodes, asking whether the channel works because the model's queries already look at the site being written. Writing the payload at *real advice text* still works (**+0.1429 [+0.1008, +0.1855]**) but is **worse** than writing it at a reserved placeholder (**−0.0373 [−0.0723, −0.0029]**), so site salience is eliminated and the original configuration is locally optimal among everything tested. Four paired views, including the arm's formatting cost: it ends **30 percentage points [+20.4, +40.4]** more episodes malformed than its own control.

Each landing page carries its arms, its per-arm rates and its headline contrasts recomputed from the campaign's episode table through the campaign's own instance-cluster bootstrap, checked against the frozen `results.json` rather than transcribed from it, and links to the other three so the ladder can be walked from any of them.

## Paper and data

- **Writeup (106pp PDF):** [writeup_rational_agents.pdf](./writeup_rational_agents.pdf) — the full program record on clean data, including the B7 contamination post-mortem and every published-vs-clean correction.
- **Public datasets:** [2026.RA.Negotiation-Campaigns](https://huggingface.co/datasets/siddharthmb/2026.RA.Negotiation-Campaigns) (frozen P1–P4 record, 55GB) · [2026.RA.Divergence-DPO-Pairs](https://huggingface.co/datasets/siddharthmb/2026.RA.Divergence-DPO-Pairs) (12,761 turn-level oracle preference pairs) · [2026.RA.Frontier-and-Scale-Cells](https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells) (frontier-API + 32B + framing + the clean `rebaseline_b2` campaign, 1.06GB). Newer campaign datasets: [2026.RA.Five-Seat-Frontier-Negotiation](https://huggingface.co/datasets/siddharthmb/2026.RA.Five-Seat-Frontier-Negotiation) · [2026.RA.Public-Oracle-Gap](https://huggingface.co/datasets/siddharthmb/2026.RA.Public-Oracle-Gap) · [2026.RA.Five-Seat-Qwen3-8B-Robustness](https://huggingface.co/datasets/siddharthmb/2026.RA.Five-Seat-Qwen3-8B-Robustness) · [2026.RA.Five-Seat-Framing-Qwen-Robustness](https://huggingface.co/datasets/siddharthmb/2026.RA.Five-Seat-Framing-Qwen-Robustness) · [2026.RA.Fairness-GRPO](https://huggingface.co/datasets/siddharthmb/2026.RA.Fairness-GRPO) (+ adapter repos) · [2026.RA.ToM-Hidden-Preference-Probe](https://huggingface.co/datasets/siddharthmb/2026.RA.ToM-Hidden-Preference-Probe) · [2026.RA.NBS-Channel-Comparison](https://huggingface.co/datasets/siddharthmb/2026.RA.NBS-Channel-Comparison) · [2026.RA.QKV-Attention-Interface](https://huggingface.co/datasets/siddharthmb/2026.RA.QKV-Attention-Interface). Note the two older datasets predate the contamination audit — screen episodes with `interlens.arena.engine.gen_failures()`, not `parse_ok`.
