AI BENCHMARK PROFILE
BigCodeBench-Hard
BigCodeBench-Hard is a subset of 148 challenging programming tasks from BigCodeBench, designed to evaluate large language models' ability to solve complex, real-world programming problems. These tasks require diverse function calls from multiple libraries across 7 domains including computation, networking, data analysis, and visualization. The benchmark tests compositional reasoning and the ability to implement complex instructions that span 139 libraries with an average of 2.8 libraries per task.
- Released
- Unknown
- Readiness
- Paper only
- Primary field
- General AI
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.