AI BENCHMARK PROFILE
SentinelBench
SentinelBench is an open-source benchmark for time-evolving monitoring tasks. It contains 100 tasks across 10 synthetic web environments (email, calendars, finance, etc.) with scripted events, measuring task completion, reaction time, and resource use.
- Released
- 2026-06-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Long-running monitoring tasks require sustained attention rather than continuous action. SentinelBench captures this class and quantifies the tradeoff between responsiveness and cost, enabling comparison of agent designs.
Motivation
AI agents are increasingly asked to carry out work that spans minutes, hours, or longer.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.