FAM-Bench
FAM-Bench evaluates multimodal language and vision-language models on Food-as-Medicine reasoning. It includes 2500 expert-verified instances across 13 health conditions, with two tasks: dish-level suitability assessment (judging if a dish is suitable for a condition) and comparative dish analysis (ranking four dishes by suitability).
- Released
- 2026-05-29
- Readiness
- Paper only
- Primary field
- Health & Life Sciences
Why it matters
The benchmark fills a gap in food AI evaluation by testing whether models can integrate ingredient, visual, and clinical nutrition constraints to make condition-aware food decisions. It provides a standardized testbed for comparing models on grounded health-aware reasoning, which is relevant for applications in nutrition guidance and chronic disease management.
Motivation
Food-as-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appropriate for a specific health condition.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.