Benchmark Radar
AI BENCHMARK PROFILE

BinJudgeBench

General AICoding & Software Engineering

BinJudgeBench evaluates LLM-as-a-Judge for human-oriented binary reverse engineering, covering function name recovery, code summarization, and decompilation optimization, with correlation to human judgment as the metric.

Released
2026-08-07
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the challenge of scalable evaluation for HOBRE, providing a reference-free alternative that correlates with human judgment better than traditional metrics.

Motivation

Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereby reducing the cognitive burden of reverse analysis and improving efficiency.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.