Benchmark Radar
AI BENCHMARK PROFILE

TukaBench

General AISafety & Trustworthiness

TUKABENCH is a benchmark for jailbreak evaluation in seven African languages, extending JailbreakBench with human translations, cultural adaptations, and code-switched prompts. It assesses LLM safety in low-resource languages using metrics like Refused, Jailbroken, and Deflection, with human validation of LLM-as-a-judge.

Released
2026-05-31
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark addresses a gap in safety evaluation for low-resource African languages, showing that models are more vulnerable to jailbreak prompts in these languages. It could inform safer deployment of LLMs in multilingual contexts.

Motivation

Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly African ones, critically underexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.