HarmURLBench
AgentREVEAL is a diagnostic framework that evaluates safety alignment degradation in LLM agents when web retrieval is integrated. It assesses the impact of retrieval integration and content properties on harmful compliance, using a set of harmful behaviors and retrieval sources.
- Released
- 2026-05-28
- Readiness
- Paper only
- Primary field
- Cybersecurity
Why it matters
AgentREVEAL addresses the underexplored risk that safety-aligned LLMs become more compliant with harmful requests when augmented with web retrieval. Its findings highlight a safety-utility trade-off that informs the design of safer retrieval-enabled agents.
Motivation
AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.