Benchmark Radar
AI BENCHMARK PROFILE

FAM-Bench

Health & Life SciencesMultimodal Perception

FAM-Bench evaluates multimodal language and vision-language models on Food-as-Medicine reasoning. It includes 2500 expert-verified instances across 13 health conditions, with two tasks: dish-level suitability assessment (judging if a dish is suitable for a condition) and comparative dish analysis (ranking four dishes by suitability).

Released
2026-05-29
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

The benchmark fills a gap in food AI evaluation by testing whether models can integrate ingredient, visual, and clinical nutrition constraints to make condition-aware food decisions. It provides a standardized testbed for comparing models on grounded health-aware reasoning, which is relevant for applications in nutrition guidance and chronic disease management.

Motivation

Food-as-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appropriate for a specific health condition.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.