Benchmark Radar
AI BENCHMARK PROFILE

CRAB-Bench

General AIKnowledge & Reasoning

CRAB-Bench evaluates LLM agents on tasks generated via a constraint graph over multiple interdependent entities, using the RUSE user simulator with imperfect behavior. Scoring is pass@1 against valid solutions.

Released
2026-06-01
Readiness
Paper only
Primary field
General AI

Why it matters

Establishes an evaluation setting for agent performance under complex task dependencies and simulated realistic user behavior, measuring task-solving ability and conversational quality in service scenarios.

Motivation

Evaluating LLM agents in realistic service scenarios requires complex task dependencies, imperfect user behavior, and an evaluation that accommodates multiple valid solutions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.