Benchmark Radar
AI BENCHMARK PROFILE

VAKRA

General AIKnowledge & ReasoningSearch & RetrievalIBM Research

Evaluates multi-hop tool-use agents across live executable APIs and retrieval, scoring via policy adherence, exact match, and groundedness judges.

Released
2026-08-12
Readiness
Runnable
Primary field
General AI

Why it matters

Enterprise agent evaluations need compositional multi-source reasoning with policy constraints; VAKRA offers a public leaderboard and reproducible harness.

Motivation

Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.