Liron Pantanowitz, MD, PhD, MHA
Matthew G. Hanna, MD
Todd R. Minnigh
August 2026—At the University of Pittsburgh’s Computational Pathology and AI Center of Excellence (CPACE), artificial intelligence research in disease detection begins with a practical reality: most of the data still lives on glass.
Academic pathology departments hold decades of archived slides documenting common and rare disease presentations across histopathology, cytology, hematology, and other slides. These collections represent longitudinal data sets that are uniquely valuable for training AI systems and education. Yet until digitized, they remain physically siloed and inaccessible to computational workflows.
Since August 2025, CPACE has been engaged in a large-scale archive digitization initiative designed specifically to convert legacy educational slides into research-ready digital assets. The effort is being conducted in collaboration with Leidos QTC Health Services, which provides the mobile high-volume scanning platform supporting the project. They refer to the system they are developing as the ScanVan.
For Liron Pantanowitz, MD, PhD, MHA, chair and professor of pathology at the University of Pittsburgh Medical Center and director of CPACE, the rationale is straightforward.
“If we want AI systems to detect cancer earlier and more accurately, we need large, diverse, well-annotated data sets,” Dr. Pantanowitz says. “Many institutions already have that data—it’s just stored on glass slides.”
Early cancer detection increasingly depends on identifying subtle morphologic patterns that may not be obvious on initial review. AI models trained on broad, diverse slide sets can learn features associated with early-stage disease, precursor lesions, and rare tumor subtypes.
Some retrospective archives are particularly powerful in this regard because they can link morphology with known outcomes.
“Archives give us something prospective workflows can’t,” says Matthew Hanna, MD, vice chair of pathology informatics and associate professor of pathology, University of Pittsburgh Medical Center, who is involved in the initiative. “They provide longitudinal context. That’s essential for training and validating detection models.”
Slides prepared over decades often exhibit mounting media bubbles, faded stains, coverslip misalignment, ink annotations, debris, or edge cracks. Preparation standards have evolved over time, adding further variability.
“These are not pristine slides coming straight from histology,” Dr. Hanna notes. “If you try to process them through a single high-throughput system without triage, your failure rates climb quickly.”
To address this, the Pittsburgh team adopted a parallel scanning architecture within the ScanVan.
The system runs five scanners in parallel: four Leica Biosystems Aperio GT 450 units and one Pramana SpectralHT-4 robotic scanner. All slides are digitized at 40× magnification to preserve the morphologic detail required for AI development and future human review.
Slides are visually assessed before scanning and routed accordingly. Structurally sound slides are directed to the Leica Biosystems units for rapid processing, while slides with visible imperfections are assigned to the more fault-tolerant platform.
Operating entirely within the ScanVan with two scanning technologists in overlapping shifts managing loading and unloading, the system exceeds 3,000 slides per day.
“Our goal is to create a process and method that others—like the VA or pathology labs that cannot afford their own digital pathology scanning facility—can leverage,” Dr. Pantanowitz says. “For us the project is as much about collecting the data as it is about refining the collection of data as slides are rapidly digitized.”
After scanning, each digital slide undergoes AI-based quality assessment. Automated checks evaluate focus, tissue coverage, scan artifacts, background consistency, and potential stitching errors. Slides flagged by the system are reviewed and rescanned if necessary.
“You can’t build reliable detection algorithms on unreliable images,” Dr. Hanna says. “Automated QC gives us an additional layer of assurance before slides enter research pipelines.”
All images are stored in DICOM format. At 40× resolution, storage requirements are significant—approximately 1 terabyte per 1,000 slides.
To streamline approvals and avoid complex integration with hospital networks, digital slides housed within the ScanVan are transferred via physical media rather than direct network connection.
The scanning infrastructure operates within a 40- × 8-foot containerized unit positioned on site. The space is fully air-conditioned, with continuous temperature monitoring to protect instruments and slide integrity and ensure technician comfort.
“The goal isn’t just digitization for its own sake,” Dr. Pantanowitz says. “It’s to create data sets from archives, anywhere and anytime, that allow us to detect disease earlier and with greater confidence.”
Rare cancers may represent one of the most compelling applications for archive digitization. While individual institutions often encounter only a limited number of these cases, digitized archives make it possible to identify and assemble cohorts at a scale previously difficult to achieve.
According to the National Cancer Institute, more than 500 cancer types are considered rare, yet collectively they account for approximately 25 percent of all cancer diagnoses in the United States. Despite this prevalence, many remain understudied because sufficiently large data sets are difficult to assemble.
“Rare cancers are exactly where archive digitization can have an outsized impact,” Dr. Hanna says. “No single institution may have enough cases to support robust model development, but multiple archives allow us to identify and assemble cohorts that would otherwise be impossible to study.”
For computational pathology researchers, digitized archives transform decades of historical material into searchable data sets that can support education, biomarker discovery, AI development, and pharmaceutical research. By bringing scanning infrastructure to the archive rather than moving large slide collections offsite, institutions can accelerate digitization while maintaining control of valuable pathology assets.
The Pittsburgh experience demonstrates that with deliberate and portable workflow design, parallel architectures, and AI-assisted quality control, archival slides can be transformed into digital infrastructure capable of supporting the next generation of cancer detection research.
Dr. Pantanowitz is chair and professor of pathology and Dr. Hanna is vice chair of pathology informatics and associate professor—both in the Department of Pathology, University of Pittsburgh Medical Center. Todd Minnigh is chief of medical imaging innovation at Leidos.