A Framework for Quantitative Evaluation and Causal Provenance Analysis of Hallucinations in Large Language Models

Authors

  • Hao-En Lee University of California, Irvine, Irvine, USA Author

DOI:

https://doi.org/10.70088/razew369

Keywords:

large language models, hallucination evaluation, causal provenance, evidence perturbation, retrieval augmentation

Abstract

Hallucination evaluation requires a clear distinction between unsupported content, detection errors, and the underlying mechanisms that generate an answer. This paper presents a comprehensive framework linking response-level measurement, annotation provenance, and controlled evidence perturbations. An executable instantiation analyzes the public RAGTruth corpus, comprising 17,790 archived responses associated with 2,965 distinct sources. Four lightweight detectors are rigorously evaluated using the official source-disjoint split, with uncertainty estimated by source-cluster bootstrap resampling. The corpus contains 7,664 responses with annotated hallucinations and 14,289 labeled spans. Excluding spans marked as implicitly true reduces the positive response count to 7,010, demonstrating that the annotation policy fundamentally changes the evaluation target. On 2,700 test responses, combined lexical and metadata logistic regression achieves an area under the receiver operating characteristic curve of 0.873, compared with 0.842 for metadata alone. Task-matched evidence replacement increases the lexical random forest score by 0.294 on average. Empty evidence instead decreases the lexical logistic regression score by 0.416, exposing a critical extrapolation failure. These perturbations identify the sensitivity of the detector to evidence while keeping responses fixed. They leave generation-level causal effects unestimated. Ultimately, the proposed framework provides an auditable evaluation procedure and establishes a practical boundary between measured detector behavior and hypotheses about hallucination origins.

Downloads

Published

2026-10-03