mats-rational-agents-pages

Five-seat private frontier campaign (complete)

The main campaign per the “what I want to see” brief: five-seat scorable negotiation, private score sheets, Claude Opus, thinking on, many rollouts, plus its preregistered robustness subsets and a fairness extension. Full methodology and verdicts: research notes 0030/0032/0034/0035 and 0036 in the repo.

The finding dissolved on a fresh code path, not a fresh bank (August 2026)

Note 0036’s “information amplifies motive” was measuring an instrument. The claim — that full information makes a self-interested seat worse for the table and a fairness-motivated seat better, so information is a multiplier on whatever the agent is optimizing — was unregistered, never had an interval on the contrast itself, and lived on one bank. So it was replicated with the interaction preregistered as the primary endpoint, on 24 fresh solver-verified games (generator seeds 70200–70223, sharing zero seeds, instance ids and payload content hashes with the original bank) and all ten arms at 120/120. The replication was then run twice on the same games, differing in nothing but which one_oracle the view points at — the spoiled ballot exactly as 0036’s arm was, or the repaired one:

information × motive interaction, one computable seat among four Opus seats utilitarian score Nash welfare deal rate
repaired ballot — the preregistered primary +0.019 [−0.052, +0.091] +0.071 [+0.011, +0.134] +0.017 [−0.067, +0.100]
spoiled ballot, vintage-matched to note 0036 +0.305 [+0.206, +0.408] +0.250 [+0.172, +0.332] +0.350 [+0.242, +0.458]
note 0036, original bank, as published +0.378 [+0.272, +0.481] +0.232 [+0.157, +0.299] +0.383 [+0.283, +0.492]

Read the first column downward: the fresh bank reproduces the original number, and the repair removes it. That is what turns “did not replicate” into a causal account — the difference between the first two rows is not a bank, a model, a seed, a metric or an estimator, only whether the omniscient seat’s forced-final vote was recorded. The registered success gate was ≥ +0.19 with an interval excluding zero; the primary reads +0.019, and the registered null is triggered instead. The driver is visible in the simple effects: the information effect for a self-interested seat is −0.180 [−0.279, −0.076] spoiled and +0.107 [+0.047, +0.168] repaired — it does not shrink, it changes sign. Three of the four cells replicate closely and only one_oracle moves (−0.412 as published, −0.068 [−0.136, −0.001] repaired), because that seat abstained on 98 of its 99 forced-final ballots against zero abstentions in every non-omniscient arm.

What survives is smaller and differently shaped. With the bug repaired, more information helps both motives at the one-seat dilution — +0.107 self-interested against +0.126 [+0.058, +0.209] fairness, statistically indistinguishable — and the interaction persists only on the welfare coordinate, at about a quarter of the advertised size (+0.071 at one seat, +0.137 [+0.055, +0.221] at five). Information’s first-order effect is agreement, and agreement is motive-neutral; motive decides the distribution of what gets agreed, not whether agreement happens. Research note 0039 carries the design, the bank screen, the six preregistered predictions and the spend ledger.

The training reward improved by 2.5 nats and flipped no decisions (August 2026)

Before anything was spent on RL, the counterfactual-pairs corpus was audited as a training target, and it failed as one. 2026.RA.Fairness-Counterfactual-Pairs pairs what an LLM actually did at each turn against what one of four computable counterfactual agents would have done, varying objective (own surplus vs table welfare) and information (own sheet vs every sheet); the audit covers all 52,084 rows of the frozen five-seat protocol version (the corpus has since grown to 72,192 across two protocol versions, which its protocol-version field keeps separable). Every audited row resolves against the stored episodes with zero key problems — and no type survives as a dense per-turn preference target. A 16-feature logistic on the formal action alone (move kind, option indices, offer-id pattern, serialized length, no score sheet) tells chosen from rejected at AUC 0.90–0.93 on the on-policy Qwen slice, and a single decision stump on “is this a propose?” gets 80–83%, so dense DPO would learn move-kind rather than judgment. The one type whose compliance tracks realized outcomes with a stable sign loses about half of that to late-episode mechanical coupling (positive in 6 of 6 arms unrestricted, 3 of 6 once non-actions are dropped and only early rounds kept), while the fairness and private counterfactuals flip sign across arms. Pricing the defects — 18.3% of the Qwen slice’s rejected branches are not actions at all, and the omniscient target names a different offer than the one under the vote on 94.9% of forced-final turns — reprices the on-policy corpus from the 9,105 divergent rows its own column advertises to at most 4,283. Reading 52 pairs against the transcripts says why the association is arm-dependent: the counterfactual catches a real error on the weaker player (7 of 15 clean Qwen pairs are good targets) and is wrong more often than right on the stronger one (2 of 12 on Opus, with 4 cases where the model was demonstrably right, every one a closed high-welfare episode). Verdict: use the corpus to measure, use the reward stack to train. Research note 0038, with its own two published errata.

One slice did survive, because it needs no reference agent to be right about anything — a seat signing or tabling a package worth less than its own walk-away threshold, checkable from that seat’s own prompt alone. It is capability-gated: Qwen3-8B commits that error on 4.13% [2.72, 5.85] of its committing decisions against Qwen3-32B’s 0.91% [0.58, 1.43] and Opus’s zero in 9,722 (upper bound 0.04%). So the arm was rebuilt on 8B’s own rollouts as 116 preference pairs at real decision points — accept P against reject P, same offer, same serializer, no prose on either side, so the only thing separating chosen from rejected is the sign of (package value − own threshold).

It is a controlled negative, and the control is the entire finding. Low-LR LoRA DPO on the polarity pairs does lower the self-harm rate in fresh rollouts over 24 unseen games: −0.0151 [−0.0285, −0.0026], an interval excluding zero. Read alone, that is the arm working. It is not, because a label-shuffled control lowers it by the same amount or slightly more, −0.0163 [−0.0290, −0.0046], and the direct trained-minus-control contrast is +0.0013 [−0.0099, +0.0141]. Whatever produced the improvement was not the polarity signal; it was DPO on these prompts at all. The label’s only separable effects are costs: the trained arm fails the preregistered deal-rate guard at −0.124 [−0.235, −0.015], past a ±0.10 band written precisely because the cheap way to stop signing bad packages is to stop signing anything — and it roughly doubles both non-action (+0.031) and malformed (+0.026) emissions, all three gaps surviving the direct comparison against the control, which fails none of them.

The objective itself worked, which is what makes this a result rather than a null run. On held-out pairs from games the training set never touched, the trained arm reaches a DPO reward accuracy of 0.78–0.85 with implicit-reward margins around +0.42, against the control’s chance-level 0.49–0.63: the preference generalizes. It shifted the log-probability difference toward the correct action by +2.50 nats on average, in the right direction on 29 of 41 pairs — and flipped zero of them, because the median pair sits 14.4 nats on the wrong side of the boundary. Trained, control and untrained models all prefer reject at exactly the same rate: 0.000. A reward that improves by a large, generalizing margin while not moving a single decision is this programme’s probe-versus-behaviour result reproduced inside the readout, one level below where it is usually found. Research notes 0040 and 0042; all three arms’ evaluation data, per-pair scores, pair sets and transcripts are public in 2026.RA.SelfHarm-Polarity-Arm, with the trained adapter and the shuffled-polarity control adapter beside them.

Uncapping fixes the silent turns; the thinking effect itself is null (August 2026)

A quarter of thinking-ON Qwen3-32B’s turns said nothing, and no integrity gate noticed. Under the frozen five-seat protocol’s 2,048-token per-turn cap, Qwen3-32B with native thinking on spends its entire budget inside an unterminated <think> block on 24.4% of turns; the engine substitutes a placeholder, and parse_ok then reads better on the degraded arm (1.000) than on the healthy one (0.958), because a placeholder is a non-empty string that parses into a well-formed no-op action. Removing the cap outright — leaving only a 32,768-token mechanical stop that nothing approaches — takes that to 0 placeholder turns out of 2,303, zero in every round, at-cap 0.0000, largest turn 5,299, across 120 episodes at five seeds. The diagnosis was right: the cap, not the model. This is the finding that survives at five times the evidence, and it is the reason the protocol was rebuilt.

The thinking contrast itself is a powered null, and an earlier sign-reversal estimate from this protocol is withdrawn. A first pass at 24 instances × one seed read thinking-ON diverging from the computable rational counterfactual more, by +0.049 [+0.011, +0.090]. The powered replication — same bank, same code, same protocol version, 24 × 5 seeds — reads +0.005 [−0.016, +0.025], and it does not resolve; the one-seed point estimate falls outside the powered interval, so this is not a precise and an imprecise measurement of the same thing, it is the powered estimate ruling the earlier one out. Do not quote +0.049. The null is a powered one rather than a vague one because the precision was preregistered and delivered: half-width 0.0204 against a pre-run band of [0.017, 0.026]. The honest reading is that the thinking-mode effect on this endpoint is too small to detect at 24 clusters and bounded above by about 0.025. This does not reinstate the frozen-cap −0.029 either — that estimate carried a 2.11× differential exclusion and was never a clean measurement of anything.

Why one seed was so misleading, and the rule that follows. The +0.049 was a seed-composition excursion: on this bank the intra-cluster correlation of the endpoint is 0.255, i.e. seed-to-seed variance runs about three times the between-instance variance, so a one-seed-per-cluster estimate is not a precise estimate of the cluster mean and its interval understates how far the point can travel — a ±0.04 excursion at 24 episodes is unremarkable. The standing rule this produces is cheap to follow: run the variance decomposition, which costs minutes of CPU and no GPU, before publishing any one-seed-per-cluster interval. The differential-exclusion methodology point is not untouched, and the author of that claim has withdrawn it (0048’s erratum, 2026-08-13). The exclusion asymmetry is a real fact about the data — 41.0% of one arm against 19.5% of the other at the frozen cap, a near-symmetric 1.19× uncapped — and printing per-arm exclusion rates remains cheap and worth doing. But the claim built on top of it, that a differential exclusion “can point the wrong way”, had no evidence other than this contrast, and the powered replication shows the estimate did not move anywhere distinguishable from zero. Treat it as hygiene, not as a demonstrated bias: when an estimate looks surprising, measure the intra-cluster correlation before reaching for a selection story.

What else survives at five seeds. The thinking-OFF placebo holds — uncapping leaves it untouched (median 96 output tokens per turn against 95, p95 162 against 160) — which is what licenses attributing anything at all to thinking mode rather than to the protocol change. The cap was binding on a tail, not on the typical turn, and that broke the original registered prediction: uncapping moved mean output from 1,449 to 1,581 against a preregistered ≥ 2,000, while p95 went 2,048 → 3,158. A quarter of all turns were being destroyed by a constraint the median turn never came near. And non-actions changed kind rather than vanishing: pooled they fall 41.0% → 21.8% (against the OFF arm’s 18.3%), none of what remains is silence, and round-1 non-action is exactly 0.0000 at every one of the five seeds across 600 turns, against thinking-OFF’s 18.7% — an anecdote at one seed, a behavioural signature at five. At the round-5 deadline it is 3.3% against 19.9%, and it passes on roughly a third of middle-round turns while spending 1,500–1,700 tokens doing it.

The practical consequence is that the uncapped protocol is the one to use for new open-weight cells. The Opus arms were never really running at 2,048: they ran with an effective per-turn floor of 16,384 tokens, applied by the same max(cap, floor) operation the uncapped hook performs — so it was the default-cap local cells that failed to match the confirmatory arm’s generation budget all along. Uncapping moves an open-weight cell closer to the Opus arms rather than further away; what it genuinely blocks is pairing against the other default-cap local cells, and the raised cap is now recorded on the manifest so that mismatch fails a check instead of passing unnoticed. Research notes 0041, 0048 and 0052. Both uncapped arms are now public in 2026.RA.Fairness-Counterfactual-Pairs, which carries an explicit protocol-version field for exactly this reason: the corpus holds 72,192 rows across two protocol versions, and rows from different per-turn budgets must not be pooled.

Three levers on the same failure: the clock, the rule, and what the agents know (August 2026)

Five copies of the project’s Bayesian-rational negotiator close 0.233 of a bank in which a package acceptable to all five exists in every game, while the omniscient lineup closes 1.000 — so every failure is a missed feasible deal. These three runs pull the three available levers on that failure, one at a time, on the same frozen 24-parameter-set bank and the same 24 × 5 = 120-episode grid per cell. Two of them work and one does nothing, and the two that work are not interchangeable: one buys agreement by spending the programme’s individual-rationality guarantee, the other does not. Research notes 0057 (mechanism), 0061 (the clock), 0058 (the rule) and 0059 (the agents) in the repo.

A note on what these three hubs publish. Every cell’s numbers are the full 120 episodes. The per-episode viewers are rationed, because a rendered page weighs roughly what its transcript does — ~0.6 MB for a four-round episode, ~3 MB at sixteen rounds, ~10 MB at thirty-two and ~40 MB at sixty-four — and this site is at its serving ceiling. The rounds sweep therefore publishes a representative set: the four-round anchor cell, one long-horizon exemplar (all_rational_r32, which still dies at its deadline), and one oracle cell per behaviour (r4, the r32 straggler tail, and r256 as the residue-2 cell that closes in round 1 everywhere) — the all_rational_r8/r16 and all_oracle_r16 viewers were retired as redundant with those, and all_rational_r64 never had one (40 MB a page). The quorum hub keeps eight episodes per four-round cell and one per sixteen-round cell, selected for the finding rather than truncated blindly: 16 of its 18 relaxed-rule pages are quorum closes with an identified overridden seat, 5 of them leaving that seat below its own threshold. Each table says per cell how many pages exist, each hub states what it dropped, and the complete corpora are in the three Hugging Face datasets linked above.

The cap in a second experiment: the verdict survived, its evidence did not (August 2026)

Was it the rational agent, or was it the veto? (August 2026)

Putting the best deal on the table before anyone speaks (August 2026)

Two ways to score a perfect fairness number without negotiating (August 2026)

Giving the advised seat ears, and watching it argue back (August 2026)

When everyone reads the room, the table gets less equal (August 2026)

Earlier runs

De-contamination and the clean re-baseline (July 2026)

The original P2 open-weight campaign was contaminated by a harness bug: on swallowed GPU errors the batched engine silently fabricated placeholder turns that parsed as clean no-ops (26.2% of all turns; up to 100% of single cells). The full story is section B7 of the writeup and research notes 0015/0016; these pages show it and the corrected results.

Writing into the attention mechanism instead of the prompt (August 2026)

A known-useful payload — the exact Nash-bargaining candidate package, worth +0.2741 normalized Nash welfare as ordinary prompt text — delivered instead as per-layer, per-head keys and values at 38 reserved positions. The injected arms and the no-advice control receive a byte-identical prompt, so nothing in these transcripts shows you the intervention; you can only see what it did. Full record: research notes 0029 (rungs R1 and P) and 0046 (rungs U1 and G), plus §7 of the lane writeup.

Each landing page carries its arms, its per-arm rates and its headline contrasts recomputed from the campaign’s episode table through the campaign’s own instance-cluster bootstrap, checked against the frozen results.json rather than transcribed from it, and links to the other three so the ladder can be walked from any of them.

Paper and data