Benchmark Radar
AI BENCHMARK PROFILE

StrategyBench

General AIKnowledge & ReasoningStrategyBench Team

StrategyBench is a benchmark for evaluating explicit strategy induction in large language models. It selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility.

Released
2026-08-24
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark addresses the gap in evaluating whether LLMs can explicitly abstract task rules from examples, which is important for adapting to data-scarce scenarios. It provides a systematic analysis of strategy induction across task variations and model configurations.

Motivation

As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.