Benchmark Radar
AI BENCHMARK PROFILE

SentinelBench

General AIKnowledge & Reasoning

SentinelBench is an open-source benchmark for time-evolving monitoring tasks. It contains 100 tasks across 10 synthetic web environments (email, calendars, finance, etc.) with scripted events, measuring task completion, reaction time, and resource use.

Released
2026-06-03
Readiness
Paper only
Primary field
General AI

Why it matters

Long-running monitoring tasks require sustained attention rather than continuous action. SentinelBench captures this class and quantifies the tradeoff between responsiveness and cost, enabling comparison of agent designs.

Motivation

AI agents are increasingly asked to carry out work that spans minutes, hours, or longer.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.