Benchmark Radar
AI BENCHMARK PROFILE

AGIEval

Law & GovernmentMathematics & Formal Sciences

A human-centric benchmark for evaluating foundation models on standardized exams including college entrance exams (Gaokao, SAT), law school admission tests (LSAT), math competitions, lawyer qualification tests, and civil service exams. Contains 20 tasks (18 multiple-choice, 2 cloze) designed to assess understanding, knowledge, reasoning, and calculation abilities in real-world academic and professional contexts.

Released
Unknown
Readiness
Paper only
Primary field
Law & Government

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.