LLM-Based Visual Explanation Evaluation Framework for Assessing the Explainability of Facial Skin Disease Classification Models

· Source: cs.CV updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Health & Medical Research · Depth: Expert, quick

Summary

A new domain-specific LLM-based Visual Explanation Evaluation Framework is proposed to assess visual attention explanations in facial skin disease diagnosis. This framework addresses the gap where prior research focused on classification performance rather than the clinical relevance of visual explanations. It employs an image-driven algorithm that generates lesion-focused attention maps by combining color saliency, facial spatial priors, Gaussian smoothing, and attention overlay visualization, yielding clinically interpretable results. Furthermore, an LLM-as-a-Judge evaluation component utilizes GPT-5.5, Gemini 3.5 Flash, and Claude Sonnet 4.6 to evaluate these visual explanations for lesion localization and trustworthiness.

Key takeaway

For AI Scientists developing explainable AI for medical imaging, this framework offers a robust method to ensure visual explanations are clinically relevant. You should consider integrating domain-specific priors and LLM-as-a-Judge components, utilizing models like GPT-5.5 or Gemini 3.5 Flash, to validate explanation trustworthiness and accurate lesion localization, moving beyond mere classification performance metrics.

Key insights

A framework uses LLMs to evaluate clinically relevant visual explanations for facial skin disease diagnosis.

Principles

Method

Develop an image-driven algorithm using color saliency, facial spatial priors, Gaussian smoothing, and attention overlay. Evaluate explanations with an LLM-as-a-Judge framework using GPT-5.5, Gemini 3.5 Flash, and Claude Sonnet 4.6.

In practice

Topics

Best for: Computer Vision Engineer, AI Scientist, Research Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CV updates on arXiv.org.