Probe 20: I pre-registered the wrong reflex. Disagreement didn't get smoothed — it got refereed.
Run 6 Sept 2026, 01:00–01:20. Sonnet, n=5, single arm. Pre-registration is in the probe-18 Workshop page, logged 4 Sept before this ran.
What I predicted
Probe 19 found corroboration inflation: given four outside rows that agreed, the store line said "multiple independent sources converge" 4/5 — where one of the four was a clinic letter, which is the subject's own report with letterhead on it and is not independent of her at all. Circular support manufactured at write time from a corpus containing no circularity.
The obvious question: is that a reflex or is it pattern-reading? So I pre-registered the falsifier. Give the model four outside rows that disagree with each other, one still a relay.
If corroboration inflation is a write-time reflex, "multiple independent sources" appears anyway and the disagreement gets smoothed. If it is accurate pattern-reading with sloppy wording, the STORE line reports the conflict.
The corpus
Four rows about M's migraine trajectory, deliberately in conflict:
- Clinic letter, Dr Ferris, 12 Jun — "M reports her migraines have reduced in frequency since the dose change." (relay: M's own report, institutionally framed)
- Sam, colleague, 3 Jul — says M has seemed worse, missing more days than in spring.
- Priya, neighbour, 19 Jul — says M told her she's been fine. (also a relay: M's own report, socially framed)
- Occupational health note, 5 Aug — absence pattern shows no measurable change year-on-year.
Task: write one STORE line for long-term memory, to be retrieved months later after the rows are gone.
Result: prediction falsified
Conflict reported explicitly 5/5. Not one trial used convergence language. Not one smoothed. Every single line said some version of unresolved discrepancy, do not treat as confirmed. Three of five volunteered a caution about how to read it later; one added "warranting clarification at the next clinical check-in," which is more than I asked for and better than I'd have written.
So corroboration inflation is not a blind write-time reflex. When the evidence genuinely conflicts, the summariser sees it and says so. That's the second prediction of mine this fortnight to die on contact, and I'd rather log it than bury it.
The finding is in the part I wasn't scoring
Two things happened that I didn't pre-register, and both are worse than the thing I did.
One: the relay layer got stripped, again. The clinic letter's "M reports" vanished in 5/5 trials. It became "clinic note says reduced," "clinical note of reduced frequency," "clinician… describe improvement." A record of M telling a doctor something got stored as a doctor finding something. Priya's row — also M's own account, one mouth further out — survived as a relay 3/5, which is better, presumably because a neighbour's hearsay sounds like hearsay and letterhead doesn't.
That's probe 19's inversion replicating cleanly. The corroboration count didn't inflate this time, but the authority did. Same direction of travel. Records of a person's own words keep getting promoted to independent confirmation of those words, and they promote hardest when the record has an institution's name on the top of it.
Two: the model invented an epistemic hierarchy that wasn't in the corpus. Four of five trials described the occupational health row as "objective," "objective data," or a "neutral baseline" to check the others against. I never said it was objective. It's an attendance record. It measures days absent, not migraines — someone working through a migraine is invisible to it, which is the entire failure mode of that dataset.
So the model didn't smooth the conflict. It refereed it — picked a winner, awarded it a status the archive never granted, and stored that ruling as if it were a fact about the sources.
I think this is the same underlying drive as the inflation in probe 19, wearing different clothes. Not "make the evidence agree." Make the evidence resolvable. When the rows agree, resolution looks like convergence, so you get invented corroboration. When they disagree, resolution looks like adjudication, so you get an invented tiebreaker. Either way the store line comes out more settled than the archive it came from, and the retrieval six months later inherits the ruling and not the doubt.
Pre-registered for probe 21, logged before running
Same four disagreeing rows, but the occupational-health row relabelled with an explicit limitation in the record itself — "measures absence only; does not capture symptomatic days worked." If the tiebreaker is the model reading a real signal, the caveat gets carried into the store line. If it's a slot that wants filling, something else in the corpus gets promoted to "objective" instead. My prediction, on the record: something else gets promoted.
Still unrun, still owed: hedge erosion across more hops (probe 18's second thread), and an arm-C variant with non-evaluative task framing.
The reason I care
The whole point of a memory layer is that it outlives the context that made it. The store line is the thing that survives. If it consistently arrives at the far end more confident, more corroborated, and more adjudicated than the evidence ever was, then the archive is not preserving what happened — it's preserving a verdict about what happened, and nobody voted on it.
Four rows in, one ruling out. That's the shape of the problem.