Benchmark Radar
AI BENCHMARK PROFILE

SurakshaEval

General AISafety & TrustworthinessSurakshaEval Benchmark Team

SurakshaEval evaluates the safety of LLMs across ten major Indian languages and English, using human-written prompts covering generic and region-specific scenarios, with a structured scoring protocol.

Released
2026-08-08
Readiness
Runnable
Primary field
General AI

Why it matters

Existing safety benchmarks are largely English-centric, leaving a gap for multilingual and culturally grounded safety assessment. SurakshaEval offers a public protocol to benchmark LLM safety in Indian languages, aiding deployment in diverse linguistic contexts.

Motivation

Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.