Benchmark Radar
AI BENCHMARK PROFILE

Claw-SWE-Bench

General AIKnowledge & ReasoningTokenRhythm Technologies

Claw-SWE-Bench is a multilingual SWE-bench-style benchmark and adapter protocol for comparing agent harnesses on coding tasks, with 350 GitHub issue-resolution instances across 8 languages and 43 repositories. Score is Pass@1 on patch correctness.

Released
2026-06-10
Readiness
Runnable
Primary field
General AI

Why it matters

Enables fair comparison of autonomous coding agents by treating harness and cost as first-class evaluation axes. Useful for developers and researchers building general-purpose coding agents.

Motivation

General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.