AI BENCHMARK PROFILE
TsuGO
TsuGO is a process-level reasoning benchmark for search efficiency in LLMs using Go life-and-death problems, parsing CoT into search trees and reporting diagnostics.
- Released
- 2026-08-13
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It adds search organization as a missing evaluation dimension beyond final-answer accuracy in LLM reasoning benchmarks.
Motivation
The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.