Benchmark Radar
AI BENCHMARK PROFILE

Vero

General AICoding & Software EngineeringSunblaze UCB

Vero evaluates repository-level verified code generation in Lean 4, with 43 multi-module instances, formal specifications, and proof-only and code-and-proof evaluation modes.

Released
2026-08-13
Readiness
Runnable
Primary field
General AI

Why it matters

It provides a testbed for measuring progress toward repository-scale verified software synthesis, where current agents fall short.

Motivation

AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.