Atlas H&E-TME: Scalable AI-Based Tissue Profiling Exceeds Expert Pathologist-Level Performance

Ryan Sargent
July 15, 2026

Hematoxylin and eosin (H&E) staining is the workhorse of pathology, performed on nearly every tissue sample examined for diagnosis or research. The resulting whole-slide images carry rich detail about how tissues and cells are organized, information that is especially valuable for understanding the tumor microenvironment (TME): the ecosystem of immune, stromal, and resident cells surrounding a tumor, increasingly recognized as a driver of disease progression and response to treatment.

But H&E images can contain millions of cells, and extracting all of that information accurately and at scale can be a significant challenge. Atlas H&E-TME, an application based on our Atlas family of pathology foundation models co-developed with Mayo Clinic, LMU Munich, and Charité – Universitätsmedizin Berlin, helps address that challenge. It identifies valid tissue, maps the tissue into 7 categories, and classifies each cell into 9 different classes, producing over 4,500 spatial features per image.

From the start, Atlas H&E-TME was built with a single benchmark in mind: expert-pathologist-level performance. Proving it clears that bar is technically challenging, because even experienced pathologists don't always agree on their calls for different cell types. Benchmarking an AI application against a single pathologist's labels can therefore be misleading: what looks like a model error may simply reflect normal variation between human experts. An application like Atlas H&E-TME also has to generalize beyond one small, carefully reviewed cohort to the many cancer types, tissue sites, and lab conditions seen in real-world practice. To address both issues, we evaluated Atlas H&E-TME along two complementary axes: accuracy and generalizability. 

We describe this evaluation in a paper which is, to our knowledge, the most comprehensive validation of an H&E-based tissue profiling system to date. Below, we summarize the key findings.

[Paper Figure 1] Example whole-slide image shown across the Tissue QC, Tissue Segmentation, and Cell Classification stages, illustrating how a single H&E slide is turned into cell-level outputs.

Accuracy: Meeting the Gold Standard

To assess how accurate Atlas H&E-TME's cell classifications really are, we needed a reference we could trust as ground truth. The standard approach is to have expert pathologists review the same cells and treat their calls as correct. But for certain cell types, like macrophages and plasma cells, even expert pathologists don't always agree. We wanted to hold Atlas H&E-TME to the highest standard so rather than rely on H&E review alone, we built a more rigorous reference to compare against.

On a focused cohort of 30 tissue sections across three cancer indications (colorectal carcinoma, non-small cell lung cancer, and urothelial bladder carcinoma), we built a reference from co-registered H&E/IHC images, with a five-marker panel labeling the principal cell populations of the TME. Five board-certified pathologists then reviewed each cell against these molecular markers and reached a consensus call, agreeing far more consistently than pathologists reviewing H&E alone, especially on challenging immune cells.

[Paper Figure 3] Inter-rater agreement (Krippendorff's α) among the five pathologists, comparing H&E-only annotation to IHC-informed annotation across the five evaluated cell classes. Agreement rises for every class, with the largest gains on the most ambiguous immune populations.

Measured against this IHC-informed consensus as ground truth, Atlas H&E-TME matched or exceeded the performance of expert pathologists working from H&E alone across every evaluated cell type, reaching a macro F1 score of 0.74 versus 0.71 for the pathologist average.

[Paper Figure 4] Per-class F1 scores for Atlas H&E-TME versus the pathologist average, measured against the IHC-informed consensus, with the macro F1 comparison (0.74 vs. 0.71) on the right. F1 is a standard accuracy measure that balances correctly identified cells against those missed or mislabeled, ranging from 0 to 1, with higher scores indicating better performance.

That result holds not just against the pathologist average but against the pathologists individually. Compared to each individual pathologist, Atlas H&E-TME performed best overall on macro F1.

[Paper Table 2] Per-class F1 score, Atlas H&E-TME versus five expert pathologists (P1–P5), each measured against the IHC-informed consensus.

We also evaluated the application's ability to recognize its own uncertainty. When Atlas H&E-TME was given the option to pass on its least confident calls, its performance rose to a macro F1 of 0.82, comfortably ahead of the corresponding pathologist score of 0.78. The cells the model set aside were disproportionately the ones it was most likely to have gotten wrong, confirming that this confidence signal is meaningful. In practice, this means Atlas H&E-TME can reliably deliver high-confidence outputs to end users, whether that output is a specific call or a decision to flag a cell as uncertain.

[Paper Figure 5] Per-class F1 scores for Atlas H&E-TME versus the pathologist average when both may abstain, scored on the cells each confidently kept (macro F1 0.82 vs. 0.78). The lower row shows coverage, the share of cells given a confident call.

Generalizability: Consistent at Scale

Accuracy on a focused cohort is one piece of the puzzle. The other piece is whether that accuracy holds across the morphological and technical diversity of real-world research. So we tested Atlas H&E-TME on a separate, larger cohort: 1,500+ images across 8 cancer types (including both primary tumors and their most common metastatic sites) with enough morphological subtypes to cover more than 90% of clinical cases per cancer type. The images came from over 25 laboratories and biobanks across Europe and the United States and 8 different scanners. Over 200,000 high-confidence pathologist annotations were generated across the cohort.

Unlike the accuracy study, this cohort didn’t compare Atlas H&E-TME against a separate pathologist score. Instead, high-confidence pathologist annotations served as the reference standard itself, and the numbers below show how closely Atlas H&E-TME's calls matched them. On primary sites, the macro F1 score ranged from 0.85 to 0.94 across the 8 cancer types, and carcinoma cell classification stayed at or above 0.96.

[Paper Table 3] Cell classification performance (F1 score) on primary sites, by cancer type.

The same consistency held on metastatic sites. Macro F1 ranged from 0.85 to 0.95, and carcinoma classification was at least 0.98 at every site.

[Paper Table 4] Cell classification performance (F1 score) on metastatic sites. "–" indicates a cell type not represented at that site.

In summary, performance was high and stable across cancer types and tissue contexts, with carcinoma classification consistently at or above 0.96. Together, the two validation axes demonstrate that Atlas H&E-TME performs at or above expert-pathologist-level accuracy and this performance generalizes across real world conditions.

Looking Ahead

Our evaluation was also honest about where H&E profiling remains challenging. Atlas H&E-TME performance was most variable for macrophages and plasma cells, particularly for bone and liver metastases where these immune cells can be hard to distinguish from the resident cells native to those tissues. However, this reflects a limitation of H&E staining itself, rather than a unique shortcoming of the application.

Atlas H&E-TME is under continuous development, with ongoing work to expand cancer type coverage and enrich the feature set. We recently applied Atlas H&E-TME to The Cancer Genome Atlas (TCGA) and released the results via OpenTME, an open-access dataset we will grow alongside future cancer-type coverage. Through the Research Access Program, academic researchers can also apply to run Atlas H&E-TME on their own images.

For the full methodology and results, the paper is available at https://arxiv.org/abs/2606.12346.

Related Articles

Meet Kwadjo Nyante, EMBA, MSc², CISSP, Chief Information Security Officer at Aignostics!

7.8.2025

Kwadjo and the Information Security team are responsible for ensuring all our operations remain secure, while protecting our customers, partners, and sensitive healthcare data from cyber threats and data handling risks.

Atlas H&E-TME: A Self-Service Tool for Comprehensive Analysis of the Tumor Microenvironment (TME)

7.5.2025

Aignostics recently announced early access to Atlas H&E-TME. Read on to learn more about the tool and sign up for a free trial!

From Bench to Bedside: Generalizable AI Model for ADC Biomarker Evaluation in NSCLC

13.5.2025

Presented at AACR 2025. This study demonstrates the potential of AI models to address key challenges in ADC biomarker evaluation for NSCLC. The strong alignment between our model predictions and pathologist assessments demonstrates the value of our automated scoring approach.