Benchmark Radar
AI BENCHMARK PROFILE

NovelAPIBench

General AIAgentsTool CallingCoding & Software Engineering

NovelAPIBench is a dynamic benchmark for evaluating LLM tool use with novel APIs, covering knowledge components and diagnostic failure categories.

Released
2026-06-02
Readiness
Paper only
Primary field
General AI

Why it matters

It addresses the gap in evaluating models' ability to acquire new APIs, providing insights into the complementary roles of retrieval and fine-tuning.

Motivation

Large language models for code generation often need to use APIs that are absent from their pretraining data.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.