AgentAbstain
AgentAbstain is a paired-task benchmark for evaluating LLM agents' ability to abstain from acting in scenarios such as ambiguity, conflicting constraints, or tool failures. It includes 263 paired tasks across 42 sandbox environments, with a proposed pipeline for generating fresh task instances.
- Released
- 2026-07-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Agent abstention is critical for safe deployment, yet existing evaluations focus on task success. This benchmark targets the gap in measuring calibrated abstention, highlighting that abstention capability is independent of general task-solving ability.
Motivation
Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.