PEFT-Arena
PEFT-Arena benchmark jointly evaluates target-domain performance and retention of pretrained capabilities for parameter-efficient finetuning methods. It covers mathematical and medical reasoning as target domains and measures general capability retention on benchmarks including BBH, IFEval, and NQ with SFT and RLVR training settings.
- Released
- 2026-05-27
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing PEFT evaluations focus mainly on downstream accuracy, overlooking retention of pretrained abilities. This benchmark provides a stability-plasticity perspective, enabling selection of fine-tuning methods that balance task adaptation and forgetting resistance, offering a more complete assessment for practical deployment decisions.
Motivation
Parameter-efficient finetuning (PEFT) has become the standard approach for adapting large language models, yet evaluations largely emphasize downstream accuracy while overlooking the retention of pretrained capabilities.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.