PreAct-Bench
PreActBench is a benchmark for predictive monitoring in LLMs, consisting of 1,000 paired ethical and unethical action trajectories across five domains. It evaluates whether models can infer if a partial trajectory will culminate in unethical action, using the Prefix Foresight F1 metric.
- Released
- 2026-06-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Safety research often detects unethical behavior only after it occurs. Predictive monitoring enables anticipation of harm before execution, and PreActBench measures this capability across models and guardrails.
Motivation
Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action trajectories toward a given objective.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.