Faithfulness hallucinations, where large language models generate outputs unsupported10 by retrieved evidence, remain a central challenge for trustworthy AI. We present a sys-11 tematic empirical evaluation of faithfulness in retrieval-augmented generation (RAG)12 systems using two benchmark datasets, HotpotQA and HaluBench, covering both multi-13 hop reasoning and single-hop hallucination detection. We analyze three small-to-mid-sized14 (2B–8B) open-weight LLMs in combination with multiple retrieval strategies, includ-15 ing sparse, dense, and hybrid approaches, as well as score-based and rank-based fusion16 techniques, enabling a comprehensive assessment of retrieval–generation interactions. By17 disentangling retrieval and generation errors, we characterize how different pipeline com-18 ponents contribute to hallucinations in RAG systems. Our analysis provides actionable19 insights and practical evaluation protocols, highlighting the critical role of robust retrieval20 and careful system design. These findings offer a benchmarking-oriented perspective for21 developing more reliable and faithful RAG systems within evaluated model scales.

Disentangling Faithfulness Hallucinations in Retrieval-Augmented Generation: A Systematic Benchmark and Analysis

MALA, CHANDANA SREE
;
Gezici, Gizem
;
Giannotti, Fosca
2026

Abstract

Faithfulness hallucinations, where large language models generate outputs unsupported10 by retrieved evidence, remain a central challenge for trustworthy AI. We present a sys-11 tematic empirical evaluation of faithfulness in retrieval-augmented generation (RAG)12 systems using two benchmark datasets, HotpotQA and HaluBench, covering both multi-13 hop reasoning and single-hop hallucination detection. We analyze three small-to-mid-sized14 (2B–8B) open-weight LLMs in combination with multiple retrieval strategies, includ-15 ing sparse, dense, and hybrid approaches, as well as score-based and rank-based fusion16 techniques, enabling a comprehensive assessment of retrieval–generation interactions. By17 disentangling retrieval and generation errors, we characterize how different pipeline com-18 ponents contribute to hallucinations in RAG systems. Our analysis provides actionable19 insights and practical evaluation protocols, highlighting the critical role of robust retrieval20 and careful system design. These findings offer a benchmarking-oriented perspective for21 developing more reliable and faithful RAG systems within evaluated model scales.
2026
Settore INFO-01/A - Informatica
Retrieval Augmented Generation, Large Language Models, Faithfulness Hallucinations, Hallucination Mitigation, Hybrid Retrieval, Knowledge Grounding, Multi-hop Reasoning
   PNRR Partenariati Estesi - FAIR - Future artificial intelligence research.
   FAIR
   Ministero dell'università e della ricerca
   PE_0000013
File in questo prodotto:
File Dimensione Formato  
draft.pdf

accesso aperto

Descrizione: not the published version
Tipologia: Accepted version (post-print)
Licenza: Creative Commons
Dimensione 594.59 kB
Formato Adobe PDF
594.59 kB Adobe PDF

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11384/169383
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact