Benchmark Radar
AI BENCHMARK PROFILE

PreAct-Bench

General AIKnowledge & Reasoning

PreActBench is a benchmark for predictive monitoring in LLMs, consisting of 1,000 paired ethical and unethical action trajectories across five domains. It evaluates whether models can infer if a partial trajectory will culminate in unethical action, using the Prefix Foresight F1 metric.

Released
2026-06-03
Readiness
Paper only
Primary field
General AI

Why it matters

Safety research often detects unethical behavior only after it occurs. Predictive monitoring enables anticipation of harm before execution, and PreActBench measures this capability across models and guardrails.

Motivation

Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action trajectories toward a given objective.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.