AI BENCHMARK PROFILE
K-FinHallu
Evaluates hallucination detection in multi-turn Korean financial RAG dialogues, with a taxonomy based on context answerability. Includes training and test splits for fine-tuning and benchmarking detectors.
- Released
- 2026-05-28
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
Targets a high-stakes multilingual domain where existing benchmarks lack coverage. Provides a resource to improve hallucination detection and refusal behavior in financial applications.
Motivation
Large Language Models (LLMs) have advanced financial automation through Retrieval-Augmented Generation (RAG), yet hallucinations remain a critical barrier to deployment in high-stakes environments.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.