AgentHijack
Evaluates the robustness of computer use agents under common environment corruptions such as pop-ups, resolution changes, and competing applications. The benchmark introduces 9 configurable corruptions and evaluates agent performance on desktop tasks using multimodal LLM-based agents, measuring task completion rates.
- Released
- 2026-05-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Real-world execution environments are imperfect, and minor corruptions can cause significant performance degradation. This benchmark quantifies agent fragility and supports the development of more robust computer use agents.
Motivation
Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digital workflows.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.