Benchmark Radar
AI BENCHMARK PROFILE

AgentHijack

General AIAgentsAgentHijack Team

Evaluates the robustness of computer use agents under common environment corruptions such as pop-ups, resolution changes, and competing applications. The benchmark introduces 9 configurable corruptions and evaluates agent performance on desktop tasks using multimodal LLM-based agents, measuring task completion rates.

Released
2026-05-25
Readiness
Runnable
Primary field
General AI

Why it matters

Real-world execution environments are imperfect, and minor corruptions can cause significant performance degradation. This benchmark quantifies agent fragility and supports the development of more robust computer use agents.

Motivation

Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digital workflows.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.