Benchmark Radar
AI BENCHMARK PROFILE

GENEB

Health & Life SciencesKnowledge & ReasoningGENEB team

GENEB evaluates frozen representations from 40 genomic foundation models across 100 DNA classification tasks in 13 functional categories, using a unified linear probing protocol with full-data, 10-shot, and 1-shot regimes. Primary metric is Matthews correlation coefficient, with rankings at overall, category, and task levels.

Released
2026-06-03
Readiness
Runnable
Primary field
Health & Life Sciences

Why it matters

Genomic model comparisons are fragmented across incompatible protocols, making claims of superiority unreliable. GENEB provides a controlled, multi-task reference for category-aware model selection, revealing that aggregate leaderboards are unstable and scale gains are inconsistent.

Motivation

Progress in genomic foundation models is difficult to assess due to fragmented benchmarks, incompatible evaluation protocols, and task-specific reporting.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.