GENEB
GENEB evaluates frozen representations from 40 genomic foundation models across 100 DNA classification tasks in 13 functional categories, using a unified linear probing protocol with full-data, 10-shot, and 1-shot regimes. Primary metric is Matthews correlation coefficient, with rankings at overall, category, and task levels.
- Released
- 2026-06-03
- Readiness
- Runnable
- Primary field
- Health & Life Sciences
Why it matters
Genomic model comparisons are fragmented across incompatible protocols, making claims of superiority unreliable. GENEB provides a controlled, multi-task reference for category-aware model selection, revealing that aggregate leaderboards are unstable and scale gains are inconsistent.
Motivation
Progress in genomic foundation models is difficult to assess due to fragmented benchmarks, incompatible evaluation protocols, and task-specific reporting.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.