AI BENCHMARK PROFILE
ChronoState
ChronoState evaluates whether a frozen language model can compose hidden elapsed-time scalars with symbolic task state to select temporal actions, using forced-choice accuracy under direct supervision.
- Released
- 2026-08-10
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The benchmark investigates a narrow mechanism for injecting time information into LLMs, but the results do not generalize broadly and the setup is primarily for probing one architectural variant.
Motivation
Temporal decisions in language-model systems often depend on both symbolic task state and elapsed wall-clock time, such as cache expiration, job completion, quota resets, deadlines, or stale sessions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.