Benchmark Radar
AI BENCHMARK PROFILE

PostTrainBench Lite

General AIAgents

PostTrainBench Lite measures whether an agent can design and execute a full post-training strategy (data, prompts, RL recipe, and eval loop) for a pretrained base model under a constrained time budget, scored as normalized mean reward over the improvement window.

Released
Unknown
Readiness
Paper only
Primary field
General AI

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.