Benchmark Radar
AI BENCHMARK PROFILE

P3B3

General AISafety & Trustworthiness

P3B3 is an expert-curated benchmark of conversational prompts for evaluating variety bias and controllability in LLMs across European and Brazilian Portuguese.

Released
2026-06-15
Readiness
Paper only
Primary field
General AI

Why it matters

Evaluates regional variety bias in Portuguese LLMs, addressing underrepresentation and controllability gaps.

Motivation

As Large Language Models (LLMs) become embedded in everyday communication, capturing regional linguistic variation is essential for reliable and equitable language use.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.