FMG-Bench
FMG-Bench evaluates English-language Christian theological triage and pastoral guidance in LLMs. It scores whether model responses match expected behaviors for triage levels: primary doctrine, secondary doctrine, prudential, and pastoral, with safety-related escalation appropriate to each scenario. The corpus includes 120 base scenarios plus 37 perturbation variants.
- Released
- 2026-05-29
- Readiness
- Runnable
- Primary field
- Health & Life Sciences
Why it matters
Evaluates a niche but real usage area where models are asked for faith and care advice. Provides a structured scoring contract for safety-critical escalation and robustness to rephrasing, which are not covered by generic QA or safety benchmarks. Useful for developers testing models in contexts where human referral matters.
Motivation
People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.