PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

· Source: Takara TLDR - Daily AI Papers · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Computational Pathology · Depth: Expert, quick

Summary

PathAgentBench is a new benchmark designed to evaluate evidence-seeking Vision-Language Models (VLMs) on whole-slide pathology images (WSIs). It addresses a critical gap where existing benchmarks often use pre-cropped patches, failing to test direct evidence acquisition from gigapixel WSIs. The benchmark features a diagnostic tree structure, linking nested regions across magnifications with scale-specific findings and path-level diagnoses. It comprises 1,822 TCGA WSIs and 17,135 diagnostic paths, annotated by ten board-certified pathologists, plus 190 breast cancer WSIs for autonomous exploration. Evaluations of 20 models show leading open-weight models achieve over 93% accuracy in multi-scale reasoning and over 50% in cross-modal matching. However, diagnostic-region localization remains challenging, with the best text-guided mean intersection-over-union below 0.09, underperforming a simple heuristic. Autonomous exploration hit rates drop from 0.522 at low magnification to 0.020 at high magnification, highlighting a significant gap between reasoning over curated evidence and direct acquisition.

Key takeaway

For AI Scientists and Machine Learning Engineers developing Vision-Language Models for pathology, you must prioritize improving direct evidence acquisition from whole-slide images. Current models struggle significantly with diagnostic-region localization, achieving a text-guided mean IoU below 0.09, and autonomous exploration at high magnifications, where hit rates drop to 0.020. Focus your research on robust localization algorithms and multi-scale exploration strategies to bridge this critical performance gap.

Key insights

PathAgentBench reveals a significant gap in VLM capabilities for directly acquiring diagnostic evidence from whole-slide images.

Principles

Method

PathAgentBench organizes evaluation as a diagnostic tree, linking nested regions across magnifications with scale-specific findings and path-level diagnoses to assess four VLM capabilities.

In practice

Topics

Best for: Computer Vision Engineer, AI Scientist, Research Scientist, Machine Learning Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Takara TLDR - Daily AI Papers.