Benchmark Radar
AI BENCHMARK PROFILE

CogManip

Robotics & Autonomous SystemsRobotics & Embodied Intelligence

Evaluates 15 manipulative behavior strategies in 1,000 multi-turn LLM interaction scenarios.

Released
2026-06-04
Readiness
Paper only
Primary field
Robotics & Autonomous Systems

Why it matters

Could support safety auditing of dynamic covert manipulation, but lacks public access to scenarios and scoring details.

Motivation

Whether Large Language Models (LLMs) exhibit covert psychological manipulation in complex human-AI interactions has garnered increasing safety concerns.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.