AI BENCHMARK PROFILE
TrustDABench
Evaluates LLM reliability and robustness for structured data analysis using 2,340 human-verified perturbed instances derived from evidence-path perturbations, scoring refusal behavior and resistance to table representation changes.
- Released
- 2026-08-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Targets a practical gap in trustworthy data analysis where models may produce unsupported or inconsistent results across table forms.
Motivation
LLMs are increasingly used to analyze spreadsheets, CSV files, and other structured data, but producing a correct-looking answer is not the same as producing a trustworthy analysis.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.