AI BENCHMARK PROFILE
EvalPlus
A rigorous code synthesis evaluation framework that augments existing datasets with extensive test cases generated by LLM and mutation-based strategies to better assess functional correctness of generated code, including HumanEval+ with 80x more test cases
- Released
- Unknown
- Readiness
- Paper only
- Primary field
- General AI
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.