Claw-SWE-Bench
Claw-SWE-Bench is a multilingual SWE-bench-style benchmark and adapter protocol for comparing agent harnesses on coding tasks, with 350 GitHub issue-resolution instances across 8 languages and 43 repositories. Score is Pass@1 on patch correctness.
- Released
- 2026-06-10
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Enables fair comparison of autonomous coding agents by treating harness and cost as first-class evaluation axes. Useful for developers and researchers building general-purpose coding agents.
Motivation
General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.