Benchmark Radar
AI BENCHMARK PROFILE

CultureForest

General AIKnowledge & Reasoning

CultureForest is a benchmark for cultural norm grounded reasoning, featuring 5,378 examples across 8 domains and 53 countries/regions, with progressive evaluation from multiple-choice to open-ended generation.

Released
2026-06-01
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark addresses the evaluation gap between cultural knowledge and its application, offering a verifiable and attributable assessment of reasoning grounded in cultural norms.

Motivation

Existing research largely reduces cultural intelligence in LLMs to a knowledge-level problem, overlooking whether models can effectively utilize their acquired knowledge in realistic scenarios.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.