Locating AI Failures: Facts vs Reasoning
Controlled development results behind MissingProof's fact/reasoning separation — including the hypothesis that failed.
The question
When an AI claim is unsupported, can an automated system tell you which component failed — a fact, a reasoning link, or a contradiction — rather than only that something failed? And does verifying reasoning links as first-class objects actually earn anything over simply checking facts?
Method
We built a 32-case labeled benchmark spanning five fabrication patterns (causal injection, unsupported synthesis, attribution escalation, precision inflation, certainty inflation), contradictions, and — critically — trap cases: supported causal and hedged claims a verifier must pass. Twenty unsupported cases were gold-labeled with a failure locus (PREMISE / BRIDGE / CONTRADICTION) before any system ran. Predictions and refutation conditions were sealed in writing before each run, and results were accepted as filed. Two systems used the identical scoring model, so the only variable was architecture: the treatment decomposed claims into facts plus reasoning links and scored both; the ablated control scored the same facts without relation composition.
Results
Failure-locus identification (20 unsupported cases):
- Decomposed fact+reasoning verification: 0.95 accuracy · 0.940 macro-F1
- Ablated control (facts only, same model): 0.30 accuracy · 0.391 macro-F1
- Majority-class floor: 0.70 accuracy
The control failed for a structural reason: without reasoning links as scored objects, it is unable to say "the link failed" — it misattributed every reasoning fabrication to the facts, which in these cases were true. On fabricated-relation cases the decomposed system's signature was consistent: facts scoring near 1.0 with the reasoning link near 0.0.
The hypothesis that failed — and why we're publishing it
Our initial sealed prediction was that decomposition would also improve detection (supported vs unsupported) over a plain entailment baseline. On this clean, short-text benchmark it did not: the baseline scored perfectly, and the refutation conditions we had filed in advance were met. We report that as filed. What survived testing is narrower and, we think, more valuable: on clean inputs the win is not "catching more errors" — it is locating them, separating fixable from unfixable, and specifying the missing evidence. Detection performance on long, messy production-style text remains under evaluation against public benchmarks and will be reported when complete.
Limitations, stated plainly
- 32 cases is a development benchmark, not an external validation; cases were authored by us (with gold labels fixed pre-run).
- Short, clean texts; production RAG outputs are longer and noisier.
- Locus gold labels derive from case construction — each case was built to instantiate exactly one failure type.
- External-benchmark evaluation (RAGTruth subset) is in progress and unreported here.
Why this matters for practice
An answer-level score and a located failure lead to different work. "Faithfulness 0.72" produces re-reading. "Facts supported, causal link unsupported, and here is the statement your sources would need" produces a fix — routed to the right owner, or to the honest conclusion that the claim must be weakened. That difference is the product.
MissingProof