Benchmark Radar
AI BENCHMARK PROFILE

OraclePhys

General AICoding & Software Engineering

Evaluates LLM ranking of stories by inter-story drift and identification of the governing story for multi-story 2-D steel frames, with scoring by top-1 accuracy and Spearman correlation against oracle truth.

Released
2026-08-23
Readiness
Runnable
Primary field
General AI

Why it matters

Provides exact finite-element ground truth for structural mechanics, enabling controlled comparison of fine-tuning objectives without human labels or LLM judging.

Motivation

oraclephys Benchmark, datasets and code for arXiv:2608.17162 benchmark finetuning llm reinforcement-learning structural-engineering # OraclePhys An end-to-end instrument — **benchmark, dataset, training pipeline** — for studying what fine-tuning objectives install in LLMs.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.