Benchmark Radar
AI BENCHMARK PROFILE

GDPevo

General AIKnowledge & ReasoningPrism-Shadow

GDPevo evaluates agent self-evolution on real business tasks across 24 groups (240 tasks) in domains like CRM, ERP, finance, and healthcare. It uses rule hybridization to attribute test-time gains to training experience, with held-out test tasks.

Released
2026-08-04
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks lack attribution of gains to training experience and face data contamination. This benchmark provides an automated pipeline for evolving benchmark instances and measures self-evolution ability.

Motivation

Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.