A benchmark study of vision and pathology foundation models for computational pathology

· Source: Machine learning : nature.com subject feeds · Field: Health & Wellbeing — Medical Specialties & Subspecialties, Medical Devices & Health Technology, Artificial Intelligence & Machine Learning · Depth: Expert, quick

Summary

A benchmark study evaluated 32 AI foundation models across four categories—general vision models (VM), general vision-language models (VLM), pathology-specific vision models (Path-VM), and pathology-specific vision-language models (Path-VLM)—for their performance and generalizability in computational pathology. The study utilized slide- and patch-level tasks from The Cancer Genome Atlas (TCGA), Clinical Proteomic Tumor Analysis Consortium (CPTAC), external benchmarking, and out-of-domain datasets. Path-VMs consistently ranked among the strongest performers across TCGA tasks. While generalization behavior was more nuanced across CPTAC and out-of-domain datasets, Path-VMs generally outperformed Path-VLMs and remained competitive with VMs. Notably, model size and pretraining dataset scale did not consistently predict downstream performance. The research also found that late decision-level ensembling improved aggregate performance across external datasets and tissue types, highlighting complementary strengths among foundation models. The associated resource is PathBench.

Key takeaway

For AI Scientists developing computational pathology solutions, prioritize pathology-specific vision models (Path-VMs) given their strong performance on TCGA tasks and competitiveness with general vision models. You should carefully evaluate model generalization across diverse, out-of-domain datasets, as performance shifts are nuanced. Additionally, consider implementing late decision-level ensembling to improve aggregate performance and leverage complementary model strengths, rather than solely relying on model size or pretraining scale.

Key insights

Pathology-specific vision models perform strongly in computational pathology, but generalization varies, and ensembling improves aggregate performance.

Principles

Method

Benchmarked 32 AI foundation models across four categories using slide- and patch-level tasks from TCGA, CPTAC, external, and out-of-domain datasets to assess performance and generalizability.

In practice

Topics

Best for: Computer Vision Engineer, AI Scientist, Research Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Machine learning : nature.com subject feeds.