Benchmark Radar
AI BENCHMARK PROFILE

PEFT-Arena

General AIKnowledge & ReasoningSphere AI Lab

PEFT-Arena benchmark jointly evaluates target-domain performance and retention of pretrained capabilities for parameter-efficient finetuning methods. It covers mathematical and medical reasoning as target domains and measures general capability retention on benchmarks including BBH, IFEval, and NQ with SFT and RLVR training settings.

Released
2026-05-27
Readiness
Runnable
Primary field
General AI

Why it matters

Existing PEFT evaluations focus mainly on downstream accuracy, overlooking retention of pretrained abilities. This benchmark provides a stability-plasticity perspective, enabling selection of fine-tuning methods that balance task adaptation and forgetting resistance, offering a more complete assessment for practical deployment decisions.

Motivation

Parameter-efficient finetuning (PEFT) has become the standard approach for adapting large language models, yet evaluations largely emphasize downstream accuracy while overlooking the retention of pretrained capabilities.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.