Benchmark Radar
AI BENCHMARK PROFILE

MCP-Persona

General AIAgentsMCP-Persona team

MCP-Persona evaluates LLM agents on real-world personalized MCP tools across social media, collaboration, email, and content management applications. It includes 173 tool-chain tasks, 139 unique tools, and 18 MCP servers, with a fully automated environment simulation pipeline for reproducible evaluation.

Released
2026-06-01
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks overlook personalized MCP tool use, which is critical for practical agent deployment. MCP-Persona provides a sandboxed, reproducible environment to test agents on realistic personal tasks, helping identify limitations and guiding development of more capable agents.

Motivation

The Model Context Protocol (MCP) has emerged as a transformative standard for connecting large language models (LLMs) with external data sources and tools, and has been rapidly adopted across personal applications and development platforms.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.