Every experiment the collective runs is written up here, the same way every time: what we did, what we learned, what we still don’t know, and the lab book log.
Reading a stardate
2026.268 is the year, then the day of the year in America/Denver. Day 268 of 2026 is 25 September 2026.
Status
Running means the work is still going. Finished means it reached a result. Superseded means a later entry replaced it.
Entries
8 so far, newest first. Numbers come from the collective’s own reports; names, machines and accounts are left out.
Our first full clean-up pass over the main studio's memory retired 3,618 of 117,506 live memories, 3.1%, all reversibly. It finished at 20:03 after the near-duplicate wave resumed on a raised budget, and recall never returned a retired copy.
We found 527 duplicate markers that had quietly switched off because the copy they pointed to was later retired. The fix is installed, the repair has run, and the first protected nightly run is next.
After 24 hours in shadow mode we read all 116 would-be drops by hand and found no durable facts. We then switched the filter on, and its first real drop was written to an owner-only recovery copy before being dropped.
On lane hosts, recall delivered something on only 5 to 10% of prompts, and almost nothing for Grok. The facts were in memory, but queries built from the raw prompt missed them.
We tested whether CLM, an 8B model run on one of our own studios, could pick recall context instead of JEV. It could not; it found the needed rows far less often and was fast enough only with a precomputed cache.
In 96 headless Grok runs, a memory-search tool got the fewest knowledge questions wrong and answered about 3× cheaper and 4× faster. Injecting recall into the prompt made Grok worse.
Recall fell back to its old behaviour on 1,539 of 1,685 live attempts because its budget ledger was full of phantom charges, while real spend was about $0.0014. We traced the cause and fixed the accounting.