Benchmark Radar
AI BENCHMARK PROFILE

ACEBench

Finance & EconomicsAgentsTool Calling

ACEBench is a comprehensive benchmark for evaluating Large Language Models' tool usage capabilities across three primary evaluation types: Normal (basic tool usage scenarios), Special (tool usage with ambiguous or incomplete instructions), and Agent (multi-agent interactions simulating real-world dialogues). The benchmark covers 4,538 APIs across 8 major domains and 68 sub-domains including technology, finance, entertainment, society, health, culture, and environment, supporting both English and Chinese languages.

Released
Unknown
Readiness
Paper only
Primary field
Finance & Economics

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.