Benchmark Radar
AI BENCHMARK PROFILE

OSGuard

General AISafety & Trustworthiness

OSGuard is a dual-granularity benchmark suite for evaluating safety in computer-use agents, with an action-level benchmark for local guardrail decisions and a risk-augmented execution suite for end-to-end evaluation.

Released
2026-06-13
Readiness
Paper only
Primary field
General AI

Why it matters

Task success alone misses unsafe shortcuts in computer-use agents. OSGuard's dual-granularity design distinguishes local recognition of unsafe actions from full-task safety, providing a more precise diagnosis of guardrail capabilities.

Motivation

Computer-use agents are increasingly evaluated by whether they complete realistic desktop and web tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.