Benchmark Radar
AI BENCHMARK PROFILE

K-FinHallu

Finance & EconomicsMultimodal PerceptionK-FinHallu Team

Evaluates hallucination detection in multi-turn Korean financial RAG dialogues, with a taxonomy based on context answerability. Includes training and test splits for fine-tuning and benchmarking detectors.

Released
2026-05-28
Readiness
Paper only
Primary field
Finance & Economics

Why it matters

Targets a high-stakes multilingual domain where existing benchmarks lack coverage. Provides a resource to improve hallucination detection and refusal behavior in financial applications.

Motivation

Large Language Models (LLMs) have advanced financial automation through Retrieval-Augmented Generation (RAG), yet hallucinations remain a critical barrier to deployment in high-stakes environments.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.