Benchmark Radar
AI BENCHMARK PROFILE

AICompanionBench

Robotics & Autonomous SystemsSafety & Trustworthiness

AICompanionBench provides a dataset of 2,123 human-AI companion conversations with safety risk annotations to evaluate LLMs-as-judges for detecting unsafe interactions.

Released
2026-06-03
Readiness
Runnable
Primary field
Robotics & Autonomous Systems

Why it matters

Supports the development of automated safety monitoring for AI companion platforms by providing a fine-grained benchmark for risk detection.

Motivation

As AI companion platforms such as Replika and Character.AI rapidly grow, concerns about unsafe human-AI interactions have intensified.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.