AI BENCHMARK PROFILE
GDPevo
GDPevo evaluates agent self-evolution on real business tasks across 24 groups (240 tasks) in domains like CRM, ERP, finance, and healthcare. It uses rule hybridization to attribute test-time gains to training experience, with held-out test tasks.
- Released
- 2026-08-04
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing benchmarks lack attribution of gains to training experience and face data contamination. This benchmark provides an automated pipeline for evolving benchmark instances and measures self-evolution ability.
Motivation
Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.