PolyWorkBench
PolyWorkBench evaluates LLM agents on multilingual, long-horizon workplace workflows across five domains: commerce, knowledge work, legal analysis, localization, and manufacturing. It includes 67 tasks, structured scoring via Grade, executable state verification with Pytest, and LLM-as-Judge diagnostics.
- Released
- 2026-07-07
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
This benchmark fills the gap of evaluating LLM agents on tasks that combine multilinguality and long-horizon execution, providing a structured protocol to compare agent performance and identify systematic failure modes in cross-lingual scenarios.
Motivation
While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across extended trajectories.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.