Benchmark Radar
AI BENCHMARK PROFILE

PolyWorkBench

General AIKnowledge & ReasoningPolyWorkBench Team

PolyWorkBench evaluates LLM agents on multilingual, long-horizon workplace workflows across five domains: commerce, knowledge work, legal analysis, localization, and manufacturing. It includes 67 tasks, structured scoring via Grade, executable state verification with Pytest, and LLM-as-Judge diagnostics.

Released
2026-07-07
Readiness
Paper only
Primary field
General AI

Why it matters

This benchmark fills the gap of evaluating LLM agents on tasks that combine multilinguality and long-horizon execution, providing a structured protocol to compare agent performance and identify systematic failure modes in cross-lingual scenarios.

Motivation

While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across extended trajectories.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.