BanglaVeilGuard
Evaluates Bangla LLM safety across six language forms using 2,366 prompts and a held-out 354-prompt split spanning unsafe, safe, and safe-sensitive requests, with deterministic response scoring and guardrail screening.
- Released
- 2026-08-22
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the lack of Bangla-specific safety evaluation by covering script variation and code-mixing that English-centric benchmarks miss, providing a repeatable protocol for measuring attack success and guardrail recall.
Motivation
Bangla large language model (LLM) safety is difficult to evaluate with English-centric or standard-script benchmarks because Bangla users routinely write across scripts, spellings, code-mixed forms, and regional registers.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.