TradeVerse
TradeVerse evaluates LLMs on longitudinal political trade negotiation understanding using reconstructed minutes of 1170 WTO meetings across 5 groups and 89 product groups, with three tasks: predicting HS codes, identifying responding countries, and generating final statements.
- Released
- 2026-08-06
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It addresses the evaluation gap of LLMs on longitudinal, multi-turn negotiation data, which is common in real-world political and institutional contexts, and provides a challenging test for tracking context and reasoning over extended interactions.
Motivation
LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.