AI BENCHMARK PROFILE
FinEvo-Bench
FinEvo-Bench evaluates self-evolving agents on 120 real-case-grounded tasks across 20 business scenes in six financial domains, with institution-provided procedures and rubrics for quality and compliance.
- Released
- 2026-08-06
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
Most agent benchmarks treat tasks independently and cannot measure learning from experience. FinEvo-Bench provides a longitudinal evaluation that measures both professional performance and self-evolution ability, filling a gap in agent benchmarking.
Motivation
Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.