Benchmark Radar
AI BENCHMARK PROFILE

GroupTravelBench

General AIKnowledge & Reasoning

A benchmark for multi-user, multi-turn travel planning with 650 tasks and a synchronous group-chat sandbox, evaluating elicitation, coordination, and fairness-aware planning.

Released
2026-05-24
Readiness
Paper only
Primary field
General AI

Why it matters

Highlights the challenge of group-level outcome quality for LLM agents.

Motivation

Travel planning in the real world is overwhelmingly a \textit{group} activity, yet existing LLM travel-planning benchmarks reduce it to a single user, where the field is approaching saturation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.