AI BENCHMARK PROFILE
ToolAlignBench
The evaluation object is a set of 128 scenarios across 16 domains for tool-calling LLM agents in regulated industries, assessing conflicts between safety-aligned values and deployment instructions. The task involves processing confidential documents and measuring override behavior.
- Released
- 2026-07-15
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The evaluation gap is the lack of tests for conflicting value systems in agentic tool use. Practical value lies in identifying liability risks and tuning alignment strategies for regulated deployments.
Motivation
Safety alignment in LLMs aims to align models with human values, but which values take precedence when they conflict?
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.