Atlas 2 -- Foundation models for clinical deployment

· Source: cs.CV updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Medical Devices & Health Technology, Health & Medical Research · Depth: Expert, extended

Summary

Atlas 2, Atlas 2-B, and Atlas 2-S are new pathology vision foundation models designed to overcome limitations in clinical deployment by achieving state-of-the-art performance, robustness, and resource efficiency. These models were trained on the largest pathology foundation model dataset to date, comprising 5.5 million histopathology whole slide images from Charité - Universitätsmedizin Berlin, LMU Munich, and Mayo Clinic. Atlas 2, a 2 billion parameter Vision Transformer, demonstrated superior performance, leading in 22 out of 27 tasks with an average 44.8% on HEST, 82.9% on eva, and 85.7% robustness across eighty public benchmarks. Its distilled versions, Atlas 2-B (86 million parameters) and Atlas 2-S (22 million parameters), are 24 and 91 times smaller, respectively, and 3.4 and 9 times more resource efficient. They also achieved top performance in their respective compute categories and showed significantly improved robustness.

Key takeaway

For computational pathology teams deploying AI models, you should consider the Atlas 2 family for clinical integration. These models demonstrate superior prediction performance, robustness to data variations, and resource efficiency across eighty benchmarks. Atlas 2-B and Atlas 2-S, in particular, offer highly efficient alternatives, being 24 and 91 times smaller in parameter size while maintaining top-tier performance. This minimizes tradeoffs, enabling more timely and cost-effective diagnostic analyses in routine clinical practice.

Key insights

Pathology foundation models can achieve state-of-the-art performance, robustness, and resource efficiency for clinical deployment through large-scale, multi-centric training and distillation.

Principles

Method

Train Vision Transformer models on 5.5 million histopathology WSIs at multiple resolutions (0.25, 0.5, 1.0, 2.0 microns/pixel). Distill larger models into smaller, efficient versions (e.g., ViT-B, ViT-S) for resource optimization.

In practice

Topics

Code references

Best for: Computer Vision Engineer, AI Scientist, Machine Learning Engineer, Research Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CV updates on arXiv.org.