Benchmark Radar
AI BENCHMARK PROFILE

TsuGO

General AIKnowledge & Reasoning

TsuGO is a process-level reasoning benchmark for search efficiency in LLMs using Go life-and-death problems, parsing CoT into search trees and reporting diagnostics.

Released
2026-08-13
Readiness
Paper only
Primary field
General AI

Why it matters

It adds search organization as a missing evaluation dimension beyond final-answer accuracy in LLM reasoning benchmarks.

Motivation

The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.