Benchmark Radar
AI BENCHMARK PROFILE

BekchiAI-Benchmark

General AIKnowledge & Reasoning

BekchiAI-Benchmark evaluates LLM agent skills using 2,057 deterministic tool-using ReAct tasks across 7 categories, scoring accuracy plus tool-call adherence, URL hallucination, source-match, and token cost.

Released
2026-08-27
Readiness
Paper only
Primary field
General AI

Why it matters

Accuracy-only leaderboards miss agent-specific failures; this benchmark adds behavioral metrics and verifier-checked answers, giving teams a repeatable way to compare and debug multi-step agent behavior.

Motivation

Large language model agents reason, call tools, and act autonomously over many steps, but their agentic skills-correctly sequencing tools, planning under dependencies, judging untrusted inputs, and grounding generated arguments-are hard to measure with accuracy-only leaderboards.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.