The PAR dataset: Prostate biopsy whole slide images from an underrepresented Middle Eastern population
Summary
The PAR dataset, a new public resource, provides 1,017 whole slide images (WSIs) derived from 339 prostate core needle biopsies of 185 patients in Erbil, Iraq. This dataset aims to address the critical lack of diverse histopathology data, as existing public datasets predominantly represent Western populations, hindering the generalizability of artificial intelligence (AI) models in digital pathology. Each slide is associated with Gleason scores and International Society of Urological Pathology grades, independently assigned by three pathologists. The images were digitized using a combination of high-throughput (Leica, Hamamatsu) and compact (Grundium) scanners. All data is de-identified and available in native formats via the BioImage Archive (accession S-BIAD2323), supporting research into grading concordance, color normalization, and cross-scanner robustness.
Key takeaway
For AI scientists and machine learning engineers developing digital pathology models, you should integrate the PAR dataset into your validation pipelines. This dataset, representing an underrepresented Middle Eastern population, directly addresses the critical need for diverse data to ensure your models generalize effectively beyond Western populations. Utilize it to test model robustness against different scanner types and to improve diagnostic accuracy across varied demographic groups.
Key insights
The PAR dataset offers diverse prostate biopsy WSIs to improve AI generalizability in digital pathology for underrepresented populations.
Principles
- AI model generalizability requires diverse population data.
- Public datasets are scarce, especially for non-Western regions.
- Multi-scanner data supports robustness evaluations.
Method
Digitized 339 prostate core needle biopsy glass slides from 185 patients into 1,017 WSIs. Associated slides with independent Gleason and ISUP grades from three pathologists. Scanned using Leica, Hamamatsu, and Grundium.
In practice
- Evaluate AI models on diverse populations.
- Analyze grading concordance across pathologists.
- Test AI robustness across different scanners.
Topics
- Digital Pathology
- Prostate Cancer
- Whole Slide Imaging
- Dataset Diversity
- AI Model Generalizability
- Gleason Score
Best for: Computer Vision Engineer, AI Scientist, Machine Learning Engineer, Research Scientist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CV updates on arXiv.org.