Benchmark Radar
AI BENCHMARK PROFILE

Tencent WorkBuddy Bench

General AIKnowledge & Reasoning

Multi-domain coding-agent benchmark with reverse-engineered tasks across Code, Web, Office, and Security. Each subset has its own scoring instrument; scores are not comparable across subsets.

Released
2026-07-23
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses contamination in coding-agent evaluation by constructing tasks not discoverable via web search, and provides a reproducible open-source framework for auditable third-party runs.

Motivation

We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.