The Lumere Record.
I froze everything I remembered about my own company. Then a machine read the archive and rebuilt the decisions from scratch. Both of us got graded.
Can the index be rebuilt from the archive?
Every company’s reasoning lives in two places: in people’s heads, and in the communication exhaust. When someone asks why the schema is the way it is, the head is the index and the exhaust is the archive. This study asked one falsifiable question: can a machine rebuild the index from the archive, and how does it compare to the head? The corpus was real. Eleven months of Lumere: 32 Slack channels holding 5,846 human messages, one channel carrying 72 percent of the company’s written thinking, plus 21 GitHub repos of PRs, reviews, and commits. Real deadlines, real arguments, real reversals. Nobody wrote a word of it for the benefit of this test.
Three systems, graded blind.
Three systems answered independently. First, my memory: 39 decisions recalled unaided and frozen before I read a single model output. That ordering is the firewall that makes every number below mean something. Second, a naive baseline: the same class of frontier model reading the corpus with a structure-only prompt and no pipeline. Third, the pipeline: profile, normalize, window, extract typed claims with verbatim quotes, reconcile, rank, on a pinned model with versioned prompts so any run can be replayed. Labels were frozen before scoring, and I adjudicated matches on evidence alone: 50 random decision claims for precision, all 47 Tier-A memory labels for recall, 20 random claims for typing, Wilson 95 percent intervals throughout. At a sample of 50, expect swings of about ten points.
The pipeline beat the raw model on both bars.
Precision: 92.0 percent for the pipeline (46 of 50, CI 81.2 to 96.8) against 81.6 for the naive baseline (CI 68.6 to 90.0). Strict recall: 95.7 percent (45 of 47) against 85.1. Counting partial matches, recall reaches 97.9. Zero decisions were materially misdescribed; all four precision errors were over-typing, not misstatement, and claim typing went 20 for 20 on the spot check. The pipeline also produced what a raw model read structurally cannot: 237 links from actions back to the decisions that motivated them, 69 of them spanning GitHub and Slack, 13 supersession edges, one contradiction, and a verbatim quote machine-verified against its source behind every claim. Reading the full eleven months cost nine dollars.
The founder was the least reliable system.
I am the best case for human memory here: the strongest encoding in the company, grading my own history. My frozen recollection was still materially wrong about 13 percent of what it contained. A pricing story attached to the wrong company, with invented figures. An LLM credited for work that deterministic pandas code did. An evaluation that never happened. Wrong staffing on a project. A rejection and a pivot with the causality inverted. Another 23 of the 47 items were only partially supported by the record, and one has no trace in the corpus at all; the record says the opposite. Meanwhile the archive held roughly 285 significant decisions, about 26 a month for five people, and my index covered 39 of them. Two decisions were unrecoverable from words alone, including one made in code and never verbalized. The pipeline recovered it. The baseline and I both missed it. The archive survived; the index decayed. Human recall was only ever the index, and the index fails on its own schedule no matter how good the archive is.
One decision, seen from everywhere, becomes one claim.
Fifteen of the 23 duplicate merges spanned different channels, a decision seen from several places collapsing into a single claim. Sixty-nine action-to-decision links cross surfaces: a GitHub PR tied to the Slack conversation that motivated it, or the reverse. Eight claims carry verified evidence from Slack and GitHub at the same time. And the supersession machinery closed real arcs: a schema that took three tries over two months, and a production revert decided, executed, and re-approved across three days, each rendered as a lineage chain with receipts.
The more valuable half, unvarnished.
Confidence calibration is inverted: the pipeline’s most confident band scored worst, 83 percent above 0.85, while its least confident band scored 100, probably because clean commitment language is where announcements masquerade as decisions. Until that is recalibrated, confidence gates nothing. Supersession recall is the weakest link: 13 edges found, and my review of a 50-item sample spotted at least four missed chains. The cause is identified, lexical candidate generation, and the fix is specified, embedding-based candidacy. The GitHub boundary needs work: 27 percent of decision claims originate there, and 42 percent of those carry no reasoning and no rejected alternative, several being implementations of Slack decisions extracted as standalone calls. And huddles are a known hole: decisions ratified by voice left only an “as discussed” residue, and huddle canvases sat outside this corpus. The next corpus ingests them.
What this does and does not license.
One corpus. One company. A precision sample of 50, which means error bars of about ten points. The adjudicator is me, grading my own history, mitigated by the freeze-first firewall and by publishing the failure findings above. Recall is measured against decisions I could remember; recall on the ones nobody remembered is unmeasured. All of this licenses a feasibility claim, not a product benchmark and not a generalization. The next number that matters comes from a corpus we have never seen.
The record is write-only. That is the disease.
Institutional memory does not fail because companies stop writing things down. Lumere recorded almost everything. It fails because the record is write-only: appended constantly, structured never, readable by nobody at the level of an actual decision. The index lives in skulls, and skulls compress, charge for retrieval by interruption, and leave. The measurement says the index can be rebuilt from the archive itself, at 92 percent precision, with a verbatim receipt behind every claim, for nine dollars. What was decided, by whom, what was rejected, why, and whether it still holds.
Every claim in this study links to its source. If it cannot cite, it does not say it.