AI BENCHMARK PROFILE
AFT-Bench
Holds task, backend, initial state, injected failure, agent, and language model fixed while varying the tool interface to measure callability versus operability.
- Released
- 2026-08-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It isolates interface-level causes of agent failures and enables controlled experiments on how tool APIs affect safe continuation.
Motivation
A tool call can be perfectly valid yet still leave an autonomous agent unable to determine what to do next.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.