StrategyBench
StrategyBench is a benchmark for evaluating explicit strategy induction in large language models. It selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility.
- Released
- 2026-08-24
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The benchmark addresses the gap in evaluating whether LLMs can explicitly abstract task rules from examples, which is important for adapting to data-scarce scenarios. It provides a systematic analysis of strategy induction across task variations and model configurations.
Motivation
As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.