AI BENCHMARK PROFILE
BinJudgeBench
BinJudgeBench evaluates LLM-as-a-Judge for human-oriented binary reverse engineering, covering function name recovery, code summarization, and decompilation optimization, with correlation to human judgment as the metric.
- Released
- 2026-08-07
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the challenge of scalable evaluation for HOBRE, providing a reference-free alternative that correlates with human judgment better than traditional metrics.
Motivation
Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereby reducing the cognitive burden of reverse analysis and improving efficiency.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.