AI BENCHMARK PROFILE
VAKRA
Evaluates multi-hop tool-use agents across live executable APIs and retrieval, scoring via policy adherence, exact match, and groundedness judges.
- Released
- 2026-08-12
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Enterprise agent evaluations need compositional multi-source reasoning with policy constraints; VAKRA offers a public leaderboard and reproducible harness.
Motivation
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.