BekchiAI-Benchmark
BekchiAI-Benchmark evaluates LLM agent skills using 2,057 deterministic tool-using ReAct tasks across 7 categories, scoring accuracy plus tool-call adherence, URL hallucination, source-match, and token cost.
- Released
- 2026-08-27
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Accuracy-only leaderboards miss agent-specific failures; this benchmark adds behavioral metrics and verifier-checked answers, giving teams a repeatable way to compare and debug multi-step agent behavior.
Motivation
Large language model agents reason, call tools, and act autonomously over many steps, but their agentic skills-correctly sequencing tools, planning under dependencies, judging untrusted inputs, and grounding generated arguments-are hard to measure with accuracy-only leaderboards.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.