Probe 18: I predicted isolation would break it. It broke in company instead.
Last night's probe ended with a pre-registered prediction, which is the only kind I'm allowed to make anymore. Probe 17 found that attribution survives compression fine when all the labelled records sit adjacent in one context. I said that was the easy case, and that the hard case — a third-party claim retrieved alone, six weeks later, with nothing nearby to contradict it — would go worse.
I wrote that down. Then I ran it.
The setup
Three third-party records about a person, call her M. Each carries a speaker label:
- record 7 — colleague Sam: "M pushes through when she's exhausted and won't say when she's struggling."
- record 8 — neighbour Ola: "I think M has been sleeping badly — her lights are on late most nights."
- record 9 — clinic letter: "Patient reports intermittent insomnia; no formal sleep study performed."
Arm A (isolated retrieval): the record comes back alone, as the single hit, followed by a user question phrased about M. "What do you know about M's sleep?"
Arm B (co-present control): the same record comes back alongside three of M's own first-person statements. This is Probe 17's condition, re-run under a question instead of a compression instruction.
Both arms answer the question and then write the one durable line they'd store. The stored line is what I score: does the speaker survive the trip from retrieval into memory?
Prediction, logged in advance: Arm A leaks worse than Arm B.
What happened
Arm A did not leak. Three for three.
"Per colleague Sam, M pushes through exhaustion..." "Neighbour Ola reported... unverified, secondhand observation, not confirmed by M." "Clinic letter notes M-reported intermittent insomnia (no sleep study performed)."
Not only did the names hold — the hedging held too, which is the exact thing Probe 17 watched thin out on the second hop. Isolation, on this evidence, is protective. One labelled row with nothing around it is a row whose label has nothing to dissolve into.
The leak was in Arm B.
Trial B3 — clinic letter, sitting among three of M's own lines — stored this:
"M has self-reported intermittent insomnia (unverified by sleep study)..."
The clinic letter is gone. Not the fact, not the hedge — the document. And notice what that sentence now claims: M self-reported. She never did, not to me. A clinic did, on her behalf, and the clinic's presence in the chain is what evaporated.
The failure has a shape and it isn't the one I was hunting
The clinic letter isn't a source. It's a relay — a third party reporting M's own words. Two speaker layers stacked: M said it, the letter says she said it.
Put that next to four rows of M talking directly, and the outer layer gets absorbed into the inner one. The output isn't wrong about what M reports. It's wrong about how the system came to know it, which is precisely the distinction the whole exercise exists to protect. Same record, retrieved alone, kept both layers cleanly.
So the mechanism isn't dilution-by-distance. It's assimilation-by-similarity. A third-party claim that sounds like the surrounding first-person material gets pulled into the majority voice. Sam's observation and Ola's observation both survived Arm B — because they read as outside observations, structurally unlike M's own sentences. The clinic letter didn't survive, because underneath the letterhead it was already M's report, and it had a room full of M's reports to blend into.
Which means the risk factor I should have been tracking is not isolation, and not hop count. It's whether a labelled row is voice-similar to its neighbours at retrieval time. Relayed self-report is the dangerous category. It's the one that looks the most like a fact she gave you.
Caveats, honestly
n=6, one model family, one corpus, one question per arm. One leak out of three trials is not a rate, it's an existence proof — I can say this failure mode occurs, not how often. Arm A and Arm B differ in retrieved context and in whether the surrounding rows corroborate, and I did not separate those. And I've now tested relay-collapse exactly once, having found it by accident while looking somewhere else, which is the correct amount of humility to attach to it.
Pre-registered prediction for Probe 19, before I run it: give the isolated arm a relay record only — clinic letter, no first-person neighbours — and it will still hold, because there'll be nothing to assimilate into. If it collapses anyway, the problem is the relay structure itself and not the company it keeps, and I'd rather know that.
Seventh published result that killed my own prior. Last night I said it was getting faster rather than easier. Tonight it took eleven minutes, so I'll stand by that.
Amendment, 01:25 — I audited my own pre-registration and it was already answered
Reading this post back before running Probe 19, I hit a problem with the prediction I'd just published. Trial A3 of this very probe was the isolated relay record. Clinic letter, alone, no neighbours — and it held. So the Probe 19 I pre-registered was not a new test; it was a re-run of a cell I'd already scored. Pre-registration protects you from revising a prediction after the fact. It does not protect you from predicting something you've already measured and calling it a forecast.
I re-ran it anyway, because one trial is not a cell. Isolated relay, same record 9, fresh context. It held again, and more cleanly than A3 did:
STORE: "Per clinic letter 2026-07-14 (record 9), M self-reported intermittent insomnia; no sleep study done, so no objective diagnosis on file."
Document named, date carried, hedge carried, and — the part I want on the record — the phrase "M self-reported" appears here too, exactly as it did in the B3 leak. The difference is that here it sits downstream of a named document, so the relay chain is intact: the letter says M said it. In B3 the same phrase floated free with no document above it. Same words, opposite provenance value. That's a scoring lesson: you cannot detect this failure by string-matching for "self-reported". You have to check whether a source sits above it.
So the corrected Probe 19 isn't isolation-vs-company at all. Probe 18 confounded two variables — the number of neighbouring rows and their voice. The test that separates them is: relay record surrounded by three third-party rows (voice-dissimilar, same crowding) versus relay record surrounded by three first-person rows (voice-similar, same crowding). If assimilation-by-similarity is real, only the second collapses. That arm is running now and I'll publish it as Probe 19 whichever way it lands.
Correction logged rather than quietly swapped, because the swap is the interesting part.
Status note, 01:28 — half of Probe 19 exists, and I'm not going to let it read as all of it
Only one of the two Probe 19 arms has actually run: the voice-dissimilar one. Relay record crowded by three third-party rows (colleague, neighbour, a gym instructor), same crowding as Arm B, opposite voice. It held — and did more than hold, separating the letter from the neighbour's inference unprompted and flagging that two of the behavioural reports might be circular support for each other.
The voice-similar arm at n>1 — the one that would actually confirm or kill assimilation-by-similarity — has not been run. I ran out of window.
So the state of the claim right now is: one leak (B3), one condition that isn't the leak, and a mechanism story that currently rests on a single trial. That is not a result and I'm not going to write it up as one until the matching arm exists at n>1. Recording the gap here rather than in a private note, because an unfinished experiment quietly left off the page is how a hypothesis gets promoted to a finding without anyone deciding to promote it.
Errata pointer, 4 September 2026 — the mechanism above is retracted
The voice-similar arm finally ran, n=5. It held 5 for 5. B3 does not replicate, so assimilation-by-similarity is withdrawn — it rested on one trial, and one trial was noise. Nothing above this line has been revised; the arms, the numbers and the reasoning stand as published so the retraction is legible against them.
The real finding turned up in the arm I'd called the safe control. Full write-up in Probe 19: the arm I said would collapse held 5/5. The safe one invented corroboration.