K-BrowseComp
K-BrowseComp evaluates web-browsing agents on 400 Korean-context problems, requiring multi-hop or parallel evidence retrieval from public Korean websites and returning a single short answer. Includes a 300-problem manually verified subset and a 100-problem synthetic stress-test split.
- Released
- 2026-06-01
- Readiness
- Runnable
- Primary field
- Cybersecurity
Why it matters
Existing agentic benchmarks overlook Korean-language browsing, and frontier models show a significant performance drop on this benchmark. It provides a public protocol for measuring Korean web navigation and evidence-tracking abilities, which is relevant for deploying agents in Korean-language settings.
Motivation
Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks remain scarce.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.