Benchmark Radar
AI BENCHMARK PROFILE

FMG-Bench

Health & Life SciencesKnowledge & Reasoning

FMG-Bench evaluates English-language Christian theological triage and pastoral guidance in LLMs. It scores whether model responses match expected behaviors for triage levels: primary doctrine, secondary doctrine, prudential, and pastoral, with safety-related escalation appropriate to each scenario. The corpus includes 120 base scenarios plus 37 perturbation variants.

Released
2026-05-29
Readiness
Runnable
Primary field
Health & Life Sciences

Why it matters

Evaluates a niche but real usage area where models are asked for faith and care advice. Provides a structured scoring contract for safety-critical escalation and robustness to rephrasing, which are not covered by generic QA or safety benchmarks. Useful for developers testing models in contexts where human referral matters.

Motivation

People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.