AI BENCHMARK PROFILE
Vero
Vero evaluates repository-level verified code generation in Lean 4, with 43 multi-module instances, formal specifications, and proof-only and code-and-proof evaluation modes.
- Released
- 2026-08-13
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It provides a testbed for measuring progress toward repository-scale verified software synthesis, where current agents fall short.
Motivation
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.