Benchmark Radar
AI BENCHMARK PROFILE

Instruct HumanEval

General AIKnowledge & Reasoning

Instruction-based variant of HumanEval benchmark for evaluating large language models' code generation capabilities with functional correctness using pass@k metric on programming problems

Released
Unknown
Readiness
Paper only
Primary field
General AI

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.