Benchmark Radar
AI BENCHMARK PROFILE

AFT-Bench

General AICoding & Software Engineering

Holds task, backend, initial state, injected failure, agent, and language model fixed while varying the tool interface to measure callability versus operability.

Released
2026-08-23
Readiness
Paper only
Primary field
General AI

Why it matters

It isolates interface-level causes of agent failures and enables controlled experiments on how tool APIs affect safe continuation.

Motivation

A tool call can be perfectly valid yet still leave an autonomous agent unable to determine what to do next.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.