Benchmark Radar
AI BENCHMARK PROFILE

TradeVerse

General AIKnowledge & Reasoning

TradeVerse evaluates LLMs on longitudinal political trade negotiation understanding using reconstructed minutes of 1170 WTO meetings across 5 groups and 89 product groups, with three tasks: predicting HS codes, identifying responding countries, and generating final statements.

Released
2026-08-06
Readiness
Paper only
Primary field
General AI

Why it matters

It addresses the evaluation gap of LLMs on longitudinal, multi-turn negotiation data, which is common in real-world political and institutional contexts, and provides a challenging test for tracking context and reasoning over extended interactions.

Motivation

LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.