Benchmark Radar
AI BENCHMARK PROFILE

BehaviorBench

General AIKnowledge & Reasoning

BehaviorBench evaluates personalized decision modeling from real-world behavioral traces, with belief and trade prediction tasks from prediction-market and on-chain records across 2,000 wallets.

Released
2026-06-01
Readiness
Paper only
Primary field
General AI

Why it matters

Provides a real-world alternative to simulated user benchmarks, testing whether personalization methods can use observed behavioral evidence effectively in decision-support settings.

Motivation

Many decision-support settings require systems that adapt to individual users, but evaluation data for this problem remain limited.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.