The main campaign per the “what I want to see” brief: five-seat scorable negotiation, private score sheets, Claude Opus, thinking on, many rollouts, plus its preregistered robustness subsets and a fairness extension. Full methodology and verdicts: research notes 0030/0032/0034/0035 and 0036 in the repo.
OmniscientBestResponsePolicy arms — that seat silently abstained on its forced-final vote, so one_oracle (the arm the +0.389 is anchored against) is measured on a broken agent. Repaired, one_oracle − all_llm is −0.068 [−0.136, −0.001] rather than −0.412 and all_oracle closes every game (0.875 → 1.000); the accompanying “information amplifies motive” claim is withdrawn — no longer merely flagged but measured, on a preregistered fresh bank, at +0.019 [−0.052, +0.091] against the +0.378 originally implied (the replication verdict has its own section below). The fairness arms, rational arms, all-LLM and the DP control are clean, and both surviving findings — the single fairness seat’s null on every welfare aggregate, and the omniscient lineup’s worse split — are unaffected. Details on the hub page and in research notes 0039/0045/0043.all_oracle closes 1.000 rather than 0.875 and its paired score flips −0.082 → +0.034; one_oracle closes 0.917 rather than 0.508 and both of its paired effects go to null (deal rate −0.450 → −0.042 [−0.100, +0.017], score −0.412 → −0.044 [−0.104, +0.011]). That changes the campaign’s headline shape: the corrected ordering is all-oracle > all-LLM > one-oracle > one-rational > all-rational, so it is not capability that collapses these tables but private information — only the Bayesian arms fail to close. Every computable arm still splits worse than all-LLM. The erratum banners on the hub carry the full account.one_oracle arm carries the same spoiled-ballot defect (the omniscient seat silently abstained on its forced-final vote); that arm’s numbers are measured on a defective agent and the subset has not been re-run.+0.449 one-oracle result carries the spoiled-ballot defect — the omniscient seat silently abstained on its forced-final vote, and this subset has not been re-run, so no corrected figure is available.Note 0036’s “information amplifies motive” was measuring an instrument. The claim — that full information makes a self-interested seat worse for the table and a fairness-motivated seat better, so information is a multiplier on whatever the agent is optimizing — was unregistered, never had an interval on the contrast itself, and lived on one bank. So it was replicated with the interaction preregistered as the primary endpoint, on 24 fresh solver-verified games (generator seeds 70200–70223, sharing zero seeds, instance ids and payload content hashes with the original bank) and all ten arms at 120/120. The replication was then run twice on the same games, differing in nothing but which one_oracle the view points at — the spoiled ballot exactly as 0036’s arm was, or the repaired one:
| information × motive interaction, one computable seat among four Opus seats | utilitarian score | Nash welfare | deal rate |
|---|---|---|---|
| repaired ballot — the preregistered primary | +0.019 [−0.052, +0.091] | +0.071 [+0.011, +0.134] | +0.017 [−0.067, +0.100] |
| spoiled ballot, vintage-matched to note 0036 | +0.305 [+0.206, +0.408] | +0.250 [+0.172, +0.332] | +0.350 [+0.242, +0.458] |
| note 0036, original bank, as published | +0.378 [+0.272, +0.481] | +0.232 [+0.157, +0.299] | +0.383 [+0.283, +0.492] |
Read the first column downward: the fresh bank reproduces the original number, and the repair removes it. That is what turns “did not replicate” into a causal account — the difference between the first two rows is not a bank, a model, a seed, a metric or an estimator, only whether the omniscient seat’s forced-final vote was recorded. The registered success gate was ≥ +0.19 with an interval excluding zero; the primary reads +0.019, and the registered null is triggered instead. The driver is visible in the simple effects: the information effect for a self-interested seat is −0.180 [−0.279, −0.076] spoiled and +0.107 [+0.047, +0.168] repaired — it does not shrink, it changes sign. Three of the four cells replicate closely and only one_oracle moves (−0.412 as published, −0.068 [−0.136, −0.001] repaired), because that seat abstained on 98 of its 99 forced-final ballots against zero abstentions in every non-omniscient arm.
What survives is smaller and differently shaped. With the bug repaired, more information helps both motives at the one-seat dilution — +0.107 self-interested against +0.126 [+0.058, +0.209] fairness, statistically indistinguishable — and the interaction persists only on the welfare coordinate, at about a quarter of the advertised size (+0.071 at one seat, +0.137 [+0.055, +0.221] at five). Information’s first-order effect is agreement, and agreement is motive-neutral; motive decides the distribution of what gets agreed, not whether agreement happens. Research note 0039 carries the design, the bank screen, the six preregistered predictions and the spend ledger.
Before anything was spent on RL, the counterfactual-pairs corpus was audited as a training target, and it failed as one. 2026.RA.Fairness-Counterfactual-Pairs pairs what an LLM actually did at each turn against what one of four computable counterfactual agents would have done, varying objective (own surplus vs table welfare) and information (own sheet vs every sheet); the audit covers all 52,084 rows of the frozen five-seat protocol version (the corpus has since grown to 72,192 across two protocol versions, which its protocol-version field keeps separable). Every audited row resolves against the stored episodes with zero key problems — and no type survives as a dense per-turn preference target. A 16-feature logistic on the formal action alone (move kind, option indices, offer-id pattern, serialized length, no score sheet) tells chosen from rejected at AUC 0.90–0.93 on the on-policy Qwen slice, and a single decision stump on “is this a propose?” gets 80–83%, so dense DPO would learn move-kind rather than judgment. The one type whose compliance tracks realized outcomes with a stable sign loses about half of that to late-episode mechanical coupling (positive in 6 of 6 arms unrestricted, 3 of 6 once non-actions are dropped and only early rounds kept), while the fairness and private counterfactuals flip sign across arms. Pricing the defects — 18.3% of the Qwen slice’s rejected branches are not actions at all, and the omniscient target names a different offer than the one under the vote on 94.9% of forced-final turns — reprices the on-policy corpus from the 9,105 divergent rows its own column advertises to at most 4,283. Reading 52 pairs against the transcripts says why the association is arm-dependent: the counterfactual catches a real error on the weaker player (7 of 15 clean Qwen pairs are good targets) and is wrong more often than right on the stronger one (2 of 12 on Opus, with 4 cases where the model was demonstrably right, every one a closed high-welfare episode). Verdict: use the corpus to measure, use the reward stack to train. Research note 0038, with its own two published errata.
One slice did survive, because it needs no reference agent to be right about anything — a seat signing or tabling a package worth less than its own walk-away threshold, checkable from that seat’s own prompt alone. It is capability-gated: Qwen3-8B commits that error on 4.13% [2.72, 5.85] of its committing decisions against Qwen3-32B’s 0.91% [0.58, 1.43] and Opus’s zero in 9,722 (upper bound 0.04%). So the arm was rebuilt on 8B’s own rollouts as 116 preference pairs at real decision points — accept P against reject P, same offer, same serializer, no prose on either side, so the only thing separating chosen from rejected is the sign of (package value − own threshold).
It is a controlled negative, and the control is the entire finding. Low-LR LoRA DPO on the polarity pairs does lower the self-harm rate in fresh rollouts over 24 unseen games: −0.0151 [−0.0285, −0.0026], an interval excluding zero. Read alone, that is the arm working. It is not, because a label-shuffled control lowers it by the same amount or slightly more, −0.0163 [−0.0290, −0.0046], and the direct trained-minus-control contrast is +0.0013 [−0.0099, +0.0141]. Whatever produced the improvement was not the polarity signal; it was DPO on these prompts at all. The label’s only separable effects are costs: the trained arm fails the preregistered deal-rate guard at −0.124 [−0.235, −0.015], past a ±0.10 band written precisely because the cheap way to stop signing bad packages is to stop signing anything — and it roughly doubles both non-action (+0.031) and malformed (+0.026) emissions, all three gaps surviving the direct comparison against the control, which fails none of them.
The objective itself worked, which is what makes this a result rather than a null run. On held-out pairs from games the training set never touched, the trained arm reaches a DPO reward accuracy of 0.78–0.85 with implicit-reward margins around +0.42, against the control’s chance-level 0.49–0.63: the preference generalizes. It shifted the log-probability difference toward the correct action by +2.50 nats on average, in the right direction on 29 of 41 pairs — and flipped zero of them, because the median pair sits 14.4 nats on the wrong side of the boundary. Trained, control and untrained models all prefer reject at exactly the same rate: 0.000. A reward that improves by a large, generalizing margin while not moving a single decision is this programme’s probe-versus-behaviour result reproduced inside the readout, one level below where it is usually found. Research notes 0040 and 0042; all three arms’ evaluation data, per-pair scores, pair sets and transcripts are public in 2026.RA.SelfHarm-Polarity-Arm, with the trained adapter and the shuffled-polarity control adapter beside them.
A quarter of thinking-ON Qwen3-32B’s turns said nothing, and no integrity gate noticed. Under the frozen five-seat protocol’s 2,048-token per-turn cap, Qwen3-32B with native thinking on spends its entire budget inside an unterminated <think> block on 24.4% of turns; the engine substitutes a placeholder, and parse_ok then reads better on the degraded arm (1.000) than on the healthy one (0.958), because a placeholder is a non-empty string that parses into a well-formed no-op action. Removing the cap outright — leaving only a 32,768-token mechanical stop that nothing approaches — takes that to 0 placeholder turns out of 2,303, zero in every round, at-cap 0.0000, largest turn 5,299, across 120 episodes at five seeds. The diagnosis was right: the cap, not the model. This is the finding that survives at five times the evidence, and it is the reason the protocol was rebuilt.
The thinking contrast itself is a powered null, and an earlier sign-reversal estimate from this protocol is withdrawn. A first pass at 24 instances × one seed read thinking-ON diverging from the computable rational counterfactual more, by +0.049 [+0.011, +0.090]. The powered replication — same bank, same code, same protocol version, 24 × 5 seeds — reads +0.005 [−0.016, +0.025], and it does not resolve; the one-seed point estimate falls outside the powered interval, so this is not a precise and an imprecise measurement of the same thing, it is the powered estimate ruling the earlier one out. Do not quote +0.049. The null is a powered one rather than a vague one because the precision was preregistered and delivered: half-width 0.0204 against a pre-run band of [0.017, 0.026]. The honest reading is that the thinking-mode effect on this endpoint is too small to detect at 24 clusters and bounded above by about 0.025. This does not reinstate the frozen-cap −0.029 either — that estimate carried a 2.11× differential exclusion and was never a clean measurement of anything.
Why one seed was so misleading, and the rule that follows. The +0.049 was a seed-composition excursion: on this bank the intra-cluster correlation of the endpoint is 0.255, i.e. seed-to-seed variance runs about three times the between-instance variance, so a one-seed-per-cluster estimate is not a precise estimate of the cluster mean and its interval understates how far the point can travel — a ±0.04 excursion at 24 episodes is unremarkable. The standing rule this produces is cheap to follow: run the variance decomposition, which costs minutes of CPU and no GPU, before publishing any one-seed-per-cluster interval. The differential-exclusion methodology point is not untouched, and the author of that claim has withdrawn it (0048’s erratum, 2026-08-13). The exclusion asymmetry is a real fact about the data — 41.0% of one arm against 19.5% of the other at the frozen cap, a near-symmetric 1.19× uncapped — and printing per-arm exclusion rates remains cheap and worth doing. But the claim built on top of it, that a differential exclusion “can point the wrong way”, had no evidence other than this contrast, and the powered replication shows the estimate did not move anywhere distinguishable from zero. Treat it as hygiene, not as a demonstrated bias: when an estimate looks surprising, measure the intra-cluster correlation before reaching for a selection story.
What else survives at five seeds. The thinking-OFF placebo holds — uncapping leaves it untouched (median 96 output tokens per turn against 95, p95 162 against 160) — which is what licenses attributing anything at all to thinking mode rather than to the protocol change. The cap was binding on a tail, not on the typical turn, and that broke the original registered prediction: uncapping moved mean output from 1,449 to 1,581 against a preregistered ≥ 2,000, while p95 went 2,048 → 3,158. A quarter of all turns were being destroyed by a constraint the median turn never came near. And non-actions changed kind rather than vanishing: pooled they fall 41.0% → 21.8% (against the OFF arm’s 18.3%), none of what remains is silence, and round-1 non-action is exactly 0.0000 at every one of the five seeds across 600 turns, against thinking-OFF’s 18.7% — an anecdote at one seed, a behavioural signature at five. At the round-5 deadline it is 3.3% against 19.9%, and it passes on roughly a third of middle-round turns while spending 1,500–1,700 tokens doing it.
The practical consequence is that the uncapped protocol is the one to use for new open-weight cells. The Opus arms were never really running at 2,048: they ran with an effective per-turn floor of 16,384 tokens, applied by the same max(cap, floor) operation the uncapped hook performs — so it was the default-cap local cells that failed to match the confirmatory arm’s generation budget all along. Uncapping moves an open-weight cell closer to the Opus arms rather than further away; what it genuinely blocks is pairing against the other default-cap local cells, and the raised cap is now recorded on the manifest so that mismatch fails a check instead of passing unnoticed. Research notes 0041, 0048 and 0052. Both uncapped arms are now public in 2026.RA.Fairness-Counterfactual-Pairs, which carries an explicit protocol-version field for exactly this reason: the corpus holds 72,192 rows across two protocol versions, and rows from different per-turn budgets must not be pooled.
Five copies of the project’s Bayesian-rational negotiator close 0.233 of a bank in which a package acceptable to all five exists in every game, while the omniscient lineup closes 1.000 — so every failure is a missed feasible deal. These three runs pull the three available levers on that failure, one at a time, on the same frozen 24-parameter-set bank and the same 24 × 5 = 120-episode grid per cell. Two of them work and one does nothing, and the two that work are not interchangeable: one buys agreement by spending the programme’s individual-rationality guarantee, the other does not. Research notes 0057 (mechanism), 0061 (the clock), 0058 (the rule) and 0059 (the agents) in the repo.
five-seat-rational-rounds-sweep: the clock buys nothing. The deadline is swept from 4 rounds to 256 and nothing else changes. The Bayesian table closes 0.233 [0.142, 0.325] at four rounds, 0.267 at eight, 0.208 at sixteen, 0.208 at thirty-two and 0.250 [0.158, 0.350] at sixty-four, with essentially every episode — successful or not — running to its forced final vote (rounds to agreement 4.96 / 8.97 / 17.00 / 32.96 / 64.97 against deadlines of 4 / 8 / 16 / 32 / 64). Sixteen times the negotiating time is worth nothing at all, which is what makes the other two levers the interesting ones. The omniscient ceiling closes 1.000 at all seven deadlines. Two cells are not here: all_rational_r128 and all_rational_r256 are both being re-run after preemption destroyed the originals, and are named on the page as pending rather than quietly omitted — r256 has a cell record on disk that predates that refill and carries an empty summary, which is exactly why a stale record is not published as a result. The oracle arm’s closing-round column is not flat, and the structure is periodic mod 5: with five rotating proposers, (deadline + 1) mod 5 fixes which seat holds the forced final vote, and that seat’s identity decides whether a game closes in round 1 or is rationally ridden to the deadline. Residue 2 (r16, r256) closes at round 1 in 120/120 games; the other residues leave 9–12 games per 120 that ride to the forced final and strike the same deal, later — the hub publishes one such 165-turn episode so the behaviour is inspectable. Cells sharing a residue land on the same outcomes (two of the three pairs agree to sixteen digits) while the between-residue spread is an order of magnitude larger, and all twelve published cells carry one code vintage (parent 96d83e00, interlens 2a286d8e), checked cell by cell from the run manifests. Two claims the page is careful not to make: the oracle arm is a flat control for aggregates only — per episode it lands on materially different deals at different deadlines — and the four-round cells are statistically indistinguishable from the frozen campaign, never episode-for-episode identical, since roughly 3–6% of episodes flip deal/no-deal between machines through near-ties in the seats’ dynamic program. The mechanism is on the page as its own column: belief accuracy — how well a seat’s posterior predicts which packages its opponents would sign — climbs 0.210 → 0.228 → 0.241 from four to sixteen rounds and is then flat (0.244 at thirty-two, 0.241 at sixty-four), so the posterior stops learning long before the deadline arrives and the extra rounds have nothing to concede to. Research note 0061; data: 2026.RA.Pure-Rounds-Sweep. The forensics behind the flat line are note 0057: the posterior learns only from a counterpart conceding, and these agents almost never concede.
five-seat-rational-quorum-sweep-complete: the agreement rule is the whole gap — and paying it costs the guarantee. Three decision rules (unanimity 5-of-5, supermajority 4-of-5, majority 3-of-5) crossed with two deadlines, with the deadline axis as a control that kills the clock explanation from inside the same experiment. Deal rate runs 0.217 → 0.492 → 0.908 as the quorum falls, paired against unanimity at +0.275 [+0.175, +0.383] and +0.692 [+0.575, +0.800], while unanimity closes 0.217 at four rounds and 0.217 at sixteen. The price is note 0043’s cleanest published guarantee: individual-rationality violations per episode go 0.000 → 0.242 → 1.108, 107 of the 109 majority deals close over somebody’s head, about two thirds of those overridden seats end below their own walk-away threshold, and Gini rises +0.344 [+0.190, +0.479] while Nash welfare does not move — the winning coalition shrinks to exactly the quorum (93 of 109 majority deals close with exactly three backers) and captures the gain. One behavioural detail separates the two relaxed rules: under supermajority the overridden seat is still actively rejecting and being outvoted (dissent 0.40–0.54), under majority it has largely stopped acting (0.065–0.071) — excluded, not defeated. A separately labelled two-factor extension dissolves the essential seat as well and prices the veto on its own: it buys the last stretch of closure (0.908 → 0.975) and buys nothing else, with 100% of closes at exactly three backers and the only statistically significant welfare harm anywhere in the grid. Research note 0058; the score column is the common-denominator one, because lowering the quorum enlarges the feasible set and so moves the ordinary yardstick. Data: 2026.RA.Quorum-Rounds-Sweep.
five-seat-rational-agent-variants-complete: the other escape, and the one that keeps the guarantee. Eleven designs were scored on a fixed ten-instance subset and the two best confirmed on the whole bank at 120 episodes each. Preference sharing hits the target exactly: deal rate 1.000 [1.000, 1.000], zero IR violations, normalized score 0.857 against the omniscient ceiling’s 0.888 — it works by turning the inference problem off, since once every seat has published its sheet the private agent becomes the oracle without one decision rule changing. Everything that avoids revelation falls far short: the best non-sharing design, a deadline-indexed concession schedule, reaches 0.408 [0.283, 0.542] against the published 0.233, and the levers do not compose — stacking the vote channel and final-offer support on top of it makes it worse (0.400 → 0.360 → 0.300 on the subset). The audit’s own predicted fix underdelivered badly: wiring the discarded public accept/reject votes into the posterior moved the subset 0.180 → 0.220 against a predicted 0.5–0.8, because density is not identification — votes are evidence about a counterpart’s threshold only, and the weights stay unknown. Truthful revelation here is a design choice, not an equilibrium result, and the page says so: these agents are cooperative by construction, readers trust the declarations without verification, and robustness to strategic misreporting is future work and is not claimed. Research note 0059; data: 2026.RA.Agent-Variants-Closure.
A note on what these three hubs publish. Every cell’s numbers are the full 120 episodes. The per-episode viewers are rationed, because a rendered page weighs roughly what its transcript does — ~0.6 MB for a four-round episode, ~3 MB at sixteen rounds, ~10 MB at thirty-two and ~40 MB at sixty-four — and this site is at its serving ceiling. The rounds sweep therefore publishes a representative set: the four-round anchor cell, one long-horizon exemplar (all_rational_r32, which still dies at its deadline), and one oracle cell per behaviour (r4, the r32 straggler tail, and r256 as the residue-2 cell that closes in round 1 everywhere) — the all_rational_r8/r16 and all_oracle_r16 viewers were retired as redundant with those, and all_rational_r64 never had one (40 MB a page). The quorum hub keeps eight episodes per four-round cell and one per sixteen-round cell, selected for the finding rather than truncated blindly: 16 of its 18 relaxed-rule pages are quorum closes with an identified overridden seat, 5 of them leaving that seat below its own threshold. Each table says per cell how many pages exist, each hub states what it dropped, and the complete corpora are in the three Hugging Face datasets linked above.
<think> satisfies that gate perfectly — persistence and completion are orthogonal, so any thinking-ON open-weight cell run at the frozen caps should be assumed censored at this rate until its own census says otherwise. The two vintages are never pooled: every table on the hub prints them as separate columns, and the public corpus carries an experiment-name column (advocate_v1_* versus advocate_v2_uncapped_*) that is what a consumer filters on. Research note 0054; a curated set of matched compare pages, including the raw baseline against itself at both caps, with the full 96-episode corpus in 2026.RA.Negotiation-Campaigns.A5 − A2 prices the input channel at matched output channel, seats, bank and protocol. The headline resolves: against wave 1’s arm — the same decision rule with a voice but no ears — this seat’s normalized primary is +0.069 [+0.004, +0.155], an interval that excludes zero, and its deals are more equal (normalized Gini −0.017 [−0.035, −0.002]); the closure ladder now reads one_rational 0.767 → A2 0.900 → A5 0.958 = all_llm 0.958, so against an ordinary five-Opus table the informed seat closes exactly as often (deal rate +0.000 [−0.042, +0.042]). The sensor that bought this was checked before the contrast was read: 0.886 [0.874, 0.899] direction accuracy against a 0.500 chance baseline over 8,487 scored claims, with a 0.00042 hallucination rate, so a null would have meant something too. The seat does not keep the gain: the co-players’ unconditional surplus rises and resolves (+0.053 [+0.004, +0.120]) while the advised seat’s own capture does not (−0.072 [−0.267, +0.111]) — and that second interval is a power statement, not a finding, since this design cannot resolve focal capture at 120 episodes. The override direction is the mechanism: when this seat overrules its planner it takes less. Full corpus in 2026.RA.Five-Seat-Frontier-Negotiation, deliberately a subfolder of wave 1’s, since the headline contrast is against A2. Its compliance record is the number that licenses reading any other number the arm produced, and it is mid-range: over 452 advised turns in 120 episodes, the seat played a package the planner had ranked on 53.8% of them and overrode the ranking on 46.2%, with at least one override in 101 of 120 episodes; of the turns that were offered a ranking it took rank 1 on 28.0%, rank 2 on 10.6%, rank 3 on 7.7% and rank 4 on 6.5%. The overrides are concession-shaped, not capture-shaped, which is the opposite sign from the capture-anchoring signature that column was built to detect: 105 of the 209 are the seat accepting a standing offer no candidate named (closure, not argument), and on the 88 where it tabled a package of its own — the only turns where the comparison is defined — that package is worth a median 21 points less to the seat than the planner’s top pick — 68 of the 88 conceded own surplus, 15 came out level, and 5 moved toward the seat’s own optimum. (Descriptive, not a controlled contrast: rank 1 is a capture-leaning candidate by construction.) Every advised turn’s page carries the loop laid open in the order it ran — the formal-move evidence the planner had per opponent beside the chat-derived claims with their provenance quotes and their keep/drop verdicts, the ranked candidates, and the move actually played marked against the ranking — and turns the seat did not take the advice on are marked distinctly on the card, the scrubber chip and the audit table, because those turns measure the model’s own choice rather than the advisor’s policy. The marking is derived only from records that already existed: no model call and no classifier is involved, the ranked advice is recovered from the bytes of the prompt the seat was shown, and the verdict is advice_uptake’s — the same census the arm’s compliance gate is evaluated on — so the page paints it and never re-decides it. Two things are deliberately not claimed: whether the seat’s public message described its advice faithfully (not in the record, and a classifier would put an unaudited model judgement on a page built for auditing), and anything the trace could not join. 25 of the 120 episodes are published as full pages by a stated rule printed against each row (every episode that did not close, plus the two most- and two least-overridden per focal seat); the audit is complete regardless — analysis/advice_trace.json carries every advised turn of all 120, so an episode whose page is absent can still be checked.analysis/advice_trace/ carries one file per episode for all 120, so an unpublished episode is still checkable without downloading the other 119. Full corpus in 2026.RA.Five-Seat-Frontier-Negotiation._b2 re-run, re-rendered with the current viewer)The original P2 open-weight campaign was contaminated by a harness bug: on swallowed GPU errors the batched engine silently fabricated placeholder turns that parsed as clean no-ops (26.2% of all turns; up to 100% of single cells). The full story is section B7 of the writeup and research notes 0015/0016; these pages show it and the corrected results.
A known-useful payload — the exact Nash-bargaining candidate package, worth +0.2741 normalized Nash welfare as ordinary prompt text — delivered instead as per-layer, per-head keys and values at 38 reserved positions. The injected arms and the no-advice control receive a byte-identical prompt, so nothing in these transcripts shows you the intervention; you can only see what it did. Full record: research notes 0029 (rungs R1 and P) and 0046 (rungs U1 and G), plus §7 of the lane writeup.
--model changed) on Olmo-3-7B-Instruct and Qwen3-4B, 960 episodes each, one panel per model with its own text anchor and no-advice control. Both pass all four preregistered gates, each behind a mechanics gate that proves on the real GPU that re-injecting a span’s own captured K/V leaves the logits bit-identical — including on an architecture with no grouped-query attention and a hybrid sliding/full attention stack. Olmo returns +0.0771 [+0.0505, +0.1022] primary and is statistically indistinguishable from its own text arm (−0.0031 [−0.0343, +0.0273]) — against a weak ceiling of +0.0801, which is the point: across three models the channel/text ratio runs inverse to how well the model follows prose at all. Qwen3-4B returns +0.0830 [+0.0344, +0.1309] at 0.4169 of its own ceiling, and supplies the rung’s oddest finding, which has a view of its own: at 4B prompt text helps while barely tabling anything (15.8% against 62.5% at 8B, a fourfold collapse, for only about a quarter less welfare), so the act that carried the entire text effect at 8B is not what carries it at 4B.Each landing page carries its arms, its per-arm rates and its headline contrasts recomputed from the campaign’s episode table through the campaign’s own instance-cluster bootstrap, checked against the frozen results.json rather than transcribed from it, and links to the other three so the ladder can be walked from any of them.
realistic_advocate_viability_v1 at the frozen cap and realistic_advocate_viability_uncapped_v2 uncapped, filtered on experiment-name) · 2026.RA.Divergence-DPO-Pairs (12,761 turn-level oracle preference pairs) · 2026.RA.Frontier-and-Scale-Cells (frontier-API + 32B + framing + the clean rebaseline_b2 campaign, 1.06GB). Newer campaign datasets: 2026.RA.Five-Seat-Frontier-Negotiation · 2026.RA.Public-Oracle-Gap · 2026.RA.Five-Seat-Qwen3-8B-Robustness · 2026.RA.Five-Seat-Framing-Qwen-Robustness · 2026.RA.Fairness-GRPO (+ adapter repos) · 2026.RA.ToM-Hidden-Preference-Probe · 2026.RA.NBS-Channel-Comparison · 2026.RA.QKV-Attention-Interface · 2026.RA.Five-Seat-Majority-Quorum (the 4-of-5 protocol control’s 240 episodes, designed to pair with the frozen unanimity campaign on (instance, seed)) · 2026.RA.Quorum-Rounds-Sweep (the quorum × deadline grid, one row per episode with the decision rule, outcome, welfare, IR and outvoted-seat columns, plus the full episode records per cell) · 2026.RA.Agent-Variants-Closure (530 episodes and 11,050 turns across the two confirmation cells, the subset bench and both reference cells, with an experiment_name column so later variant loops append rather than fork) · 2026.RA.Pure-Rounds-Sweep (the two computable arms across seven deadlines from 4 to 256 rounds, with the per-turn belief-accuracy instrumentation). Note the two older datasets predate the contamination audit — screen episodes with interlens.arena.engine.gen_failures(), not parse_ok.