AI BENCHMARK PROFILE
BigCodeBench-Full
A comprehensive benchmark that evaluates large language models' ability to solve complex, practical programming tasks via code generation. Contains 1,140 fine-grained tasks across 7 domains using function calls from 139 libraries. Challenges LLMs to invoke multiple function calls as tools and handle complex instructions for realistic software engineering and general-purpose reasoning tasks.
- Released
- Unknown
- Readiness
- Paper only
- Primary field
- General AI
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.