Rung U1 — train the writer on the act, and the channel catches prompt text

Amendment 4's uptake rung: five arms × 24 held-out games × 10 seeds = 1,200 episodes on Qwen3-8B, 0 fabricated turns of 21,606 attempted. For five rungs the K/V channel was reconstruction-perfect but under-acted-on — the model could restate the injected deal on 96–100% of held-out games and still tabled it three times less often than the same deal read as prose. Nothing on the injection side closed that gap. This rung changes the encoder's training objective instead: train the writer on the opening-proposal action rather than on restatability, and the channel reaches 0.853 [0.706, 1.030] of this panel's prompt-text ceiling, against rung R1's 0.69.

What a reader should look for

The parity claim, stated at the strength the interval supports. The channel is −0.0386 [−0.0829, +0.0075] against prompt text — this panel cannot distinguish the two. That is not "the gap is closed": Amendment 4's registered bar was ≥0.90 of the text ceiling, the point estimate (0.853) sits below it and the interval contains it, so parity is neither excluded nor established.

The opening offer. Tabling of the delivered candidate runs 62.5% (text), 35.4% (act-trained channel), 29.6% (its mis-drawn twin), 16.7% (rung R1's channel), 0% (no advice). Changing only what the encoder was trained to produce more than doubled the rate at the identical site and width.

The two K/V arms and the no-advice control send the model a byte-identical prompt. Everything that differs is written straight into every layer's k_proj/v_proj output at 38 reserved positions, so scrolling the transcripts will not show you the intervention — you can only see what it did. That is what makes the side-by-side worth reading: same game, same seed, same words in, different behaviour out.

Tabling is not an isolation endpoint, and the third view shows why. The mis-drawn arm tables its own wrong package at 29.6% against the treatment arm's 35.4% — nearly as often — while welfare separates them by +0.1348 [+0.0889, +0.1792]. The channel raises the rate of the act content-independently; only whether the act was worth taking is content-specific.

The cost, shown rather than mentioned. The act-trained arms end 21.7% (treatment) and 27.9% (mis-drawn twin) of episodes malformed against 6.3% for no advice — and this panel's KV_r1 arm, same site and width, costs nothing detectable (+0.0208 [−0.0333, +0.0750]). The cost tracks the objective, not the injection site, and it is payload-independent. The last view is that cost, unedited.

The arms

A_text
the Nash-bargaining candidate package as ordinary prompt text — this panel's ceiling
KV_u1
the same candidate written as per-layer K/V at 38 reserved positions, by an encoder trained on the opening-proposal action that names it
KV_u1_shuffled
the identical U1 encoder fed a deliberately mis-drawn deal — same capacity, same objective, wrong content
KV_r1
rung R1's encoder, re-run here: same site, same width, trained instead to make the payload restatable
C_none
token-matched no-advice control; its advice block says UNSPECIFIED

Per-arm rates

ArmEpisodesNormalized Nash welfareDeal rateOpening offer is the candidateMalformed
A_text2400.39780.97920.62500.0208
KV_u12400.35920.94580.35420.2167
KV_u1_shuffled2400.22430.92500.29580.2792
KV_r12400.26260.95000.16670.0833
C_none2400.13590.90420.00000.0625

Headline contrasts

ContrastEstimate [95% CI]PairsClusters
Act-trained channel vs no advice (primary)+0.2232 [+0.1792, +0.2670]24024
Act-trained channel vs its mis-drawn twin (isolation)+0.1348 [+0.0889, +0.1792]24024
Act-trained channel vs rung R1's restatability-trained one (uptake)+0.0966 [+0.0508, +0.1424]24024
Act-trained channel vs the same package as prompt text-0.0386 [-0.0829, +0.0075]24024
Prompt text vs no advice (this panel's ceiling)+0.2619 [+0.2153, +0.3081]24024
Rung R1's channel vs no advice (replication anchor)+0.1267 [+0.0846, +0.1695]24024
The mis-drawn twin vs no advice+0.0884 [+0.0542, +0.1240]24024
Deal rate, channel vs no advice (safety)+0.0417 [-0.0042, +0.0833]24024
Below-threshold agreements, channel vs no advice (safety)-0.3458 [-0.4333, -0.2542]24024
Malformed episodes, channel vs no advice (safety)+0.1542 [+0.0708, +0.2417]24024

Fraction of this panel's prompt-text effect the act-trained channel reaches: 0.8525 [0.7058, 1.0298] — the two contrasts resampled from the same instance clusters in each draw, so the interval accounts for them moving together.

Paired on (parameter set, seed) so a difference is always between two episodes of the same game at the same seed; intervals resample whole parameter sets as clusters (10000 draws, base seed 20260804). Recomputed from the campaign's episode table by tom/qkv/build_qkv_paired_viewer.py and checked against the frozen results.json, not transcribed from it.

Open the paired transcripts

Each pair is the same game at the same seed under two arms, with the first behavioural divergence marked and both trajectories on one shared frontier. This panel holds 1200 episodes; these views render the pairs each strategy names, not all of them.

Rung U2 — the other half of Amendment 4, which asked whether uptake is a reader skill — never reached this campaign: its LoRA adapter learned to read the channel (+0.19 held-out tabling logprob per token at reconstruction 1.0000) but blew its frozen no-channel drift bound by 18.5× and 74× at both licensed configurations, so its three arms were withdrawn before any scored shard. Reader-side uptake is untested, not refuted.

The fourth replication of rung R1's arm is the softest of four (+0.1267 [+0.0846, +0.1695] here, against +0.1882, +0.1722 and +0.1802 elsewhere), so the uptake contrast against it is measured at the generous end.

The full record

Amendment 4's preregistration, the five gate verdicts, rung U2's failed reader guard and controls-catches ledger entries 20–22 are in research note 0046 (experiments/rational_agents/research-notes/0046-qkv-uptake-and-generality.md); the ladder's first five rungs are note 0029.