AI BENCHMARK PROFILE
OraclePhys
Evaluates LLM ranking of stories by inter-story drift and identification of the governing story for multi-story 2-D steel frames, with scoring by top-1 accuracy and Spearman correlation against oracle truth.
- Released
- 2026-08-23
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides exact finite-element ground truth for structural mechanics, enabling controlled comparison of fine-tuning objectives without human labels or LLM judging.
Motivation
oraclephys Benchmark, datasets and code for arXiv:2608.17162 benchmark finetuning llm reinforcement-learning structural-engineering # OraclePhys An end-to-end instrument — **benchmark, dataset, training pipeline** — for studying what fine-tuning objectives install in LLMs.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.