AI BENCHMARK PROFILE
BTS-AgentBench
Evaluates multi-turn agent performance on 532 telemetry-derived tasks across train/dev/test splits, with additional XAI4HEAT episodes, using deterministic replayable construction and verifiable gold answers.
- Released
- 2026-08-27
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a reproducible pipeline from raw telemetry to agent benchmarks, enabling consistent evaluation of agent capabilities in industrial settings with evidence attribution and quality-gated reporting.
Motivation
Industrial sites contain large volumes of read-only telemetry, but few benchmarks specify how to compile these records into executable multi-turn agent tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.