SurakshaEval
SurakshaEval evaluates the safety of LLMs across ten major Indian languages and English, using human-written prompts covering generic and region-specific scenarios, with a structured scoring protocol.
- Released
- 2026-08-08
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing safety benchmarks are largely English-centric, leaving a gap for multilingual and culturally grounded safety assessment. SurakshaEval offers a public protocol to benchmark LLM safety in Indian languages, aiding deployment in diverse linguistic contexts.
Motivation
Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.