Benchmark Radar
AI BENCHMARK PROFILE

BanglaVeilGuard

General AISafety & Trustworthiness

Evaluates Bangla LLM safety across six language forms using 2,366 prompts and a held-out 354-prompt split spanning unsafe, safe, and safe-sensitive requests, with deterministic response scoring and guardrail screening.

Released
2026-08-22
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the lack of Bangla-specific safety evaluation by covering script variation and code-mixing that English-centric benchmarks miss, providing a repeatable protocol for measuring attack success and guardrail recall.

Motivation

Bangla large language model (LLM) safety is difficult to evaluate with English-centric or standard-script benchmarks because Bangla users routinely write across scripts, spellings, code-mixed forms, and regional registers.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.