Benchmark Radar
AI BENCHMARK PROFILE

TrustDABench

General AISafety & Trustworthiness

Evaluates LLM reliability and robustness for structured data analysis using 2,340 human-verified perturbed instances derived from evidence-path perturbations, scoring refusal behavior and resistance to table representation changes.

Released
2026-08-25
Readiness
Runnable
Primary field
General AI

Why it matters

Targets a practical gap in trustworthy data analysis where models may produce unsupported or inconsistent results across table forms.

Motivation

LLMs are increasingly used to analyze spreadsheets, CSV files, and other structured data, but producing a correct-looking answer is not the same as producing a trustworthy analysis.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.