Benchmark Radar
AI BENCHMARK PROFILE

BLUEX v2

General AIMathematics & Formal SciencesTropicAI Research

BLUEX v2 evaluates large language models on open-ended, discursive questions from the second-phase entrance exams of UNICAMP and USP (2022–2025). The dataset includes 395 questions with 919 subquestions, covering nine subjects, with 55.7% of questions containing images represented as context-aware captions. Scoring uses an LLM-as-a-judge protocol with binary rubric criteria based on official reference answers, yielding a 0–10 score.

Released
2026-06-21
Readiness
Runnable
Primary field
General AI

Why it matters

Portuguese-language evaluation of LLMs has been limited, especially for open-ended tasks requiring deep reasoning and generation. This benchmark provides a public, reusable testbed for assessing capabilities in mathematical reasoning, image understanding, and other dimensions, offering comparable scores across models for practical model selection.

Motivation

Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended, discursive tasks that demand deeper reasoning and generation capabilities.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.