AI Reasoning Error Detection
The most dangerous AI errors aren't wrong facts. They're fabricated connections between right ones.
Ask people what an AI hallucination looks like and they describe an invented fact — a fake citation, a wrong number. Those exist, and they're the easy case: fact-checking catches them. The failure class that survives review is different: every fact in the claim is individually true, and the relationship asserted between them is invented. We call this a reasoning error, and detecting it requires scoring facts and reasoning separately — which is the core of how MissingProof works.
The four reasoning-error patterns
- Causal fabrication. "Traffic doubled in April" + "the rebranding campaign ran in April" becomes "the rebranding campaign doubled traffic." Temporal adjacency laundered into causation.
- Attribution escalation. "The committee discussed the policy" becomes "the committee concluded the policy should change." Discussion is a fact; the conclusion is fabricated.
- Certainty inflation. "May reduce costs, pending the pilot" becomes "will definitely reduce costs 20%." The hedge — which was part of the source's actual claim — is stripped in the rewrite.
- Unsupported synthesis. Three true retrieved passages welded into a conclusion none of them makes. Retrieval succeeded; the synthesis is the fabrication.
Why standard detection misses this class
Similarity and aggregate faithfulness scoring compare the answer's text to the source's text. A reasoning error is constructed from the source text — so overlap is high precisely when the fabrication is worst. The measurement that catches it must ask a different question per asserted relation: is this relation stated anywhere in the sources, and separately, does it necessarily follow from the supported facts? Two facts co-existing never entails causation between them; a discussion never entails its own conclusion. Scored explicitly, these fabrications are unambiguous. Averaged into one number, they vanish.
What the separated measurement looks like
Claim: "Product B caused the company's revenue increase."
Sources establish: revenue increased · Product B sales increased · Product B is 70% of revenue.
MissingProof output: Facts ~100%. Reasoning ~0% — no source states the causal relation, and three true facts do not entail it (revenue could have risen from other lines; B's share doesn't establish B's growth as the driver). Locus: BRIDGE. Missing proof: a source statement attributing the increase to Product B.
In controlled development testing, this separation is precisely where decomposed verification outperformed a same-model control that scored facts without relation composition — the ablated control was structurally unable to say "the link failed" and misattributed every reasoning error to the facts. Details and limitations on the research page.
MissingProof