OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice

· Source: Artificial Intelligence · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Nutrition, Fitness & Lifestyle Medicine · Depth: Expert, quick

Summary

OmniFood-Bench is a new comprehensive benchmark designed to evaluate Large Vision-Language Models (VLMs) for nutrient reasoning and personalized health advice, addressing the "Systemic Information Asymmetry" between food's visual appearance and its nutritional content. Built from the MM-Food-100K dataset, it assesses VLMs across three progressive capabilities: Basic Perception (ingredients, cooking methods), Quantitative Reasoning (portion size, nutritional profiling), and Safety-Critical Advisory (disease-specific recommendations). Evaluations of six leading VLMs, including gpt-5.1, gemini-3-flash, and qwen3-vl-8B, revealed a significant "Semantic-Physical Gap." While models achieved near-human accuracy in dish identification, they exhibited catastrophic failure in mass estimation and frequently hallucinated benign yet dangerous advice for high-risk diabetic profiles. This benchmark establishes a rigorous standard for trustworthiness in autonomous agents deployed for public health.

Key takeaway

For Research Scientists developing or deploying VLMs in health-related applications, you must prioritize rigorous evaluation beyond basic perception. Your current models, including gpt-5.1 or gemini-3-flash, exhibit a "Semantic-Physical Gap." This leads to catastrophic failures in mass estimation and dangerous hallucinations for high-risk profiles. You should integrate benchmarks like OmniFood-Bench to validate quantitative reasoning and safety-critical advisory capabilities before any public health deployment.

Key insights

VLMs exhibit a "Semantic-Physical Gap," failing catastrophically in food mass estimation and safety-critical health advice despite visual recognition accuracy.

Principles

Method

OmniFood-Bench evaluates VLMs on MM-Food-100K, assessing Basic Perception, Quantitative Reasoning (mass, nutrition), and Safety-Critical Advisory (disease-specific recommendations).

In practice

Topics

Best for: CTO, VP of Engineering/Data, Director of AI/ML, AI Scientist, Research Scientist, AI Ethicist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.