Benchmark Radar
AI BENCHMARK PROFILE

LilyBench

General AIKnowledge & ReasoningCSCPadova

LilyBench evaluates symbolic music generation and understanding using LilyPond, with a 200-prompt generation suite and ten understanding tasks, scored via compile rate, descriptor similarity, and FMD.

Released
2026-06-07
Readiness
Runnable
Primary field
General AI

Why it matters

Symbolic music evaluation is fragmented; LilyBench provides a joint benchmark with public datasets and code, enabling reproducible comparisons and metric triangulation.

Motivation

Symbolic music evaluation for large language models remains fragmented across representations, datasets, and metrics.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.