Benchmark Radar
AI BENCHMARK PROFILE

TF-RefusalBench

General AIKnowledge & Reasoning

TF-RefusalBench is a multilingual benchmark for criminal-law translation and summarization derived from public Swiss Supreme Court rulings. It contains 5,200 prompts across French, German, Italian, and English, designed to trigger refusals in LLMs.

Released
2026-06-22
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark addresses the challenge of evaluating over-alignment in LLMs performing legitimate legal tasks. It provides a standardized way to measure refusal behavior across languages and task types, aiding in the selection and mitigation of models for sensitive translation and summarization work.

Motivation

While the wider applicability of LLMs in the legal field is currently debated due to their reliability and the gravity of any errors, narrow uses with well-understood and mitigated risks have emerged.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.