Loopzero

An instrument records only the failures it can represent. The rest are filed as health.

Loopzero is an independent research program on measurement validity in recursive systems — markets, recommenders, model-training loops — where a system's own output becomes its next input.

Paper: Absorbing Shadow Classes in Conjunctive Label Rules
Manuscript · August 2026 · not yet peer-reviewed

The program

The program began by asking whether recursive degradation leaves an observable footprint before overt failure. Testing that question required building a benchmark — and the benchmark turned out to be the finding.

Its label rule declared a unit degraded only when a run condition co-occurred with an eligibility condition. The eligibility clause was absorbing: once a unit crossed below the floor it could never satisfy the rule again, so the units failing most continuously were filed, by construction, as healthy controls. 69.2% of the nominal control class was in that state. The most natural early-warning statistic pointed the wrong way, and a favourable published result turned out to be an artifact of the contamination.

The program's subject is now that class of failure: what a measuring instrument is structurally unable to record, and what happens to the cases it cannot see.

The finding

3,292 of 4,755 nominal controls — 69.2% — were structurally unlabelable: fully observed, failing continuously, and incapable under the rule of ever being recorded as failures. They missed a median 15 of 20 early steps; the labeled failures missed at most 7.

A pre-registered learnability gate caught the contradiction before any detector result existed. Treating starvation as a competing terminal event and rebuilding the control class repaired the instrument: the same gate that failed now passes (TPR 0.561 against a chance rate of 0.05) — the first valid measurement the benchmark produced.

Reported in full

On the repaired instrument, the pre-registered hypothesis failed at every alarm budget (ΔTPR = −0.472, CI [−0.496, −0.444]). That null is the published result.

The protocol

Two checks, an afternoon's work, runnable on any labeled benchmark.

  • Census The shadow-class census Every eligibility condition partitions a population into labelable and unlabelable units. Report the size and behaviour of the unlabelable stratum before using the benchmark for inference.
  • Gate The learnability gate Ask whether the simplest label-derived signal separates the classes in the right direction, above chance, at a matched alarm budget. If it does not, no detector result from that population can be interpreted.

The protocol is public and will remain public. If you maintain a benchmark — public or internal — the ask is simple: agree now to run the census on it and publish the result either way. I'll help with the run. Agreeing in advance is what makes that true.

Open problems

Two questions the paper leaves open. Both are scoped, pre-registrable, and available to anyone who wants them.

  • 01 Slate-structure early warning, fair-fight edition On the purified population, does any slate-structure quantity with demonstrated depletion-independence beat the outcome-stream baseline? One indicator (δ) survived directionally; the gate and benchmark are public; the bar is the early miss rate. Candidates must show independence from frontier depletion before detector evaluation.
  • 02 The churn inversion Churn's effect reverses sign on the purified population — among clean controls only — and the paper does not explain why. Diagnose the mechanism.

Three witnesses the bridge proposed — gain, persistence, diversity. On the repaired instrument one held its predicted direction, one inverted, one went null. The bridge remains an open empirical question.

The record

Everything is checkable by a stranger.

Collaborate

The program is one person, and the next results should not be. What it needs: teams willing to run the census on instruments they maintain, researchers who want either open problem, and hostile readers who will make the claims narrower.