MissingProof

Locating AI Failures: Facts vs Reasoning

Controlled development results behind MissingProof's fact/reasoning separation — including the hypothesis that failed.

The question

When an AI claim is unsupported, can an automated system tell you which component failed — a fact, a reasoning link, or a contradiction — rather than only that something failed? And does verifying reasoning links as first-class objects actually earn anything over simply checking facts?

Method

We built a 32-case labeled benchmark spanning five fabrication patterns (causal injection, unsupported synthesis, attribution escalation, precision inflation, certainty inflation), contradictions, and — critically — trap cases: supported causal and hedged claims a verifier must pass. Twenty unsupported cases were gold-labeled with a failure locus (PREMISE / BRIDGE / CONTRADICTION) before any system ran. Predictions and refutation conditions were sealed in writing before each run, and results were accepted as filed. Two systems used the identical scoring model, so the only variable was architecture: the treatment decomposed claims into facts plus reasoning links and scored both; the ablated control scored the same facts without relation composition.

Results

Failure-locus identification (20 unsupported cases):

The control failed for a structural reason: without reasoning links as scored objects, it is unable to say "the link failed" — it misattributed every reasoning fabrication to the facts, which in these cases were true. On fabricated-relation cases the decomposed system's signature was consistent: facts scoring near 1.0 with the reasoning link near 0.0.

The hypothesis that failed — and why we're publishing it

Our initial sealed prediction was that decomposition would also improve detection (supported vs unsupported) over a plain entailment baseline. On this clean, short-text benchmark it did not: the baseline scored perfectly, and the refutation conditions we had filed in advance were met. We report that as filed. What survived testing is narrower and, we think, more valuable: on clean inputs the win is not "catching more errors" — it is locating them, separating fixable from unfixable, and specifying the missing evidence. Detection performance on long, messy production-style text remains under evaluation against public benchmarks and will be reported when complete.

Limitations, stated plainly

Why this matters for practice

An answer-level score and a located failure lead to different work. "Faithfulness 0.72" produces re-reading. "Facts supported, causal link unsupported, and here is the statement your sources would need" produces a fix — routed to the right owner, or to the honest conclusion that the claim must be weakened. That difference is the product.

See the located-failure output on your own text