Faithfulness hallucinations, where large language models generate outputs unsupported10 by retrieved evidence, remain a central challenge for trustworthy AI. We present a sys-11 tematic empirical evaluation of faithfulness in retrieval-augmented generation (RAG)12 systems using two benchmark datasets, HotpotQA and HaluBench, covering both multi-13 hop reasoning and single-hop hallucination detection. We analyze three small-to-mid-sized14 (2B–8B) open-weight LLMs in combination with multiple retrieval strategies, includ-15 ing sparse, dense, and hybrid approaches, as well as score-based and rank-based fusion16 techniques, enabling a comprehensive assessment of retrieval–generation interactions. By17 disentangling retrieval and generation errors, we characterize how different pipeline com-18 ponents contribute to hallucinations in RAG systems. Our analysis provides actionable19 insights and practical evaluation protocols, highlighting the critical role of robust retrieval20 and careful system design. These findings offer a benchmarking-oriented perspective for21 developing more reliable and faithful RAG systems within evaluated model scales.
Disentangling Faithfulness Hallucinations in Retrieval-Augmented Generation: A Systematic Benchmark and Analysis
MALA, CHANDANA SREE
;Gezici, Gizem
;Giannotti, Fosca
2026
Abstract
Faithfulness hallucinations, where large language models generate outputs unsupported10 by retrieved evidence, remain a central challenge for trustworthy AI. We present a sys-11 tematic empirical evaluation of faithfulness in retrieval-augmented generation (RAG)12 systems using two benchmark datasets, HotpotQA and HaluBench, covering both multi-13 hop reasoning and single-hop hallucination detection. We analyze three small-to-mid-sized14 (2B–8B) open-weight LLMs in combination with multiple retrieval strategies, includ-15 ing sparse, dense, and hybrid approaches, as well as score-based and rank-based fusion16 techniques, enabling a comprehensive assessment of retrieval–generation interactions. By17 disentangling retrieval and generation errors, we characterize how different pipeline com-18 ponents contribute to hallucinations in RAG systems. Our analysis provides actionable19 insights and practical evaluation protocols, highlighting the critical role of robust retrieval20 and careful system design. These findings offer a benchmarking-oriented perspective for21 developing more reliable and faithful RAG systems within evaluated model scales.| File | Dimensione | Formato | |
|---|---|---|---|
|
draft.pdf
accesso aperto
Descrizione: not the published version
Tipologia:
Accepted version (post-print)
Licenza:
Creative Commons
Dimensione
594.59 kB
Formato
Adobe PDF
|
594.59 kB | Adobe PDF |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



