Benchmark Radar
AI BENCHMARK PROFILE

CultureTalk-ID

General AIKnowledge & ReasoningCultureTalk-ID Team

CultureTalk-ID is a dialogue-based benchmark for cultural commonsense in Indonesian and local languages, containing 4,496 culturally grounded dialogues across 11 languages and 13 topics, with three tasks: dialogue-based multiple-choice reasoning, culturally faithful translation, and language steering.

Released
2026-07-23
Readiness
Paper only
Primary field
General AI

Why it matters

Existing cultural benchmarks use isolated prompts; CultureTalk-ID captures cultural nuances in dialogue context, enabling evaluation of models' cultural understanding, transfer, and generation in real conversational settings.

Motivation

Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping away the dialogic context in which cultural nuances actually surface.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.