Comparative Methods Review: Predicting Tumor Mutational Burden from Routine Pathology Slides

Authors

  • Ruilin Wen

DOI:

https://doi.org/10.61173/s3cyfa41

Keywords:

tumor mutational burden, whole-slide imaging, multiple-instance learning, self-supervised learning, foundation models

Abstract

Sequencing remains the reference standard for tumor mutational burden (TMB) but is costly, slow, and tissue intensive. Routine hematoxylin-and-eosin whole-slide images (WSIs) are inexpensive to digitize, motivating interest in whether AI can estimate TMB to support triage and prioritize sequencing. This review synthesizes more than thirty studies of TMB-from-WSI, standardizing task framing (primarily binary TMB-high versus TMB-low, with occasional regression) and evaluation practice (AUROC/AUPRC, external validation, calibration, and decision-curve analysis). Reported internal performance frequently falls around AUROC 0.70–0.82; independent-site external results are lower, approximately 0.65–0.73, yet directionally supportive. Multimodal fusion of H&E with basic clinical variables and the use of stronger representation self-supervised encoders and pathology foundation models-improve robustness, but performance remains sensitive to label definitions, class prevalence, tumor purity, and site/scanner domain shift. Reporting calibration quality, clinical net benefit, and subgroup analyses is inconsistent across studies. Overall, the current evidence supports TMB-from-WSI as a tool for triage and sequencing prioritization rather than a replacement for sequencing. This review recommends multicenter external validation, a minimal reporting set with a practical “benchmark card,” and post-deployment monitoring of discrimination, calibration, and drift. Foundation and vision–language models with few-shot adapters are promising for cross-site transfer; prospective multicenter evaluations will be pivotal for clinical credibility.

References

[1] Kather, J. N., Pearson, A. T., Halama, N., et al. (2019). Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer. Nature Medicine, 25(7), 1054–1056. https://doi.org/10.1038/s41591-019-0462-y

[2] Yamashita, R., Long, J., Longacre, T. A., et al. (2021). Deep learning model to predict microsatellite instability from routine histology of colorectal cancer: Multicentre validation. The Lancet Oncology, 22(1), 132–141. https://doi.org/10.1016/ S1470-2045(20)30535-0

[3] Echle, A., Grabsch, H. I., Quirke, P., et al. (2020). Clinicalgrade detection of microsatellite instability in colorectal tumors by deep learning of histology images. Gastroenterology, 159(4), 1406–1416.e11. https://doi.org/10.1053/j.gastro.2020.06.021

[4] Bilal, M., Raza, S. E. A., Azam, A., et al. (2021). Weakly supervised prediction of molecular pathways and key mutations in colorectal cancer from routine histology images: A retrospective study. The Lancet Digital Health, 3(12), e902– e912. https://doi.org/10.1016/S2589-7500(21)00180-1

[5] Sadhwani, A., Chang, H.-W., Behrooz, A., et al. (2021). Comparative analysis of machine learning approaches to classify tumor mutation burden in lung adenocarcinoma using histopathology images. Scientific Reports, 11, 16605. https://doi. org/10.1038/s41598-021-95747-4

[6] Jain, M. S., & Massoud, T. F. (2020). Predicting tumour mutational burden from histopathological images using multiscale deep learning. Nature Machine Intelligence, 2(6), 356–362. https://doi.org/10.1038/s42256-020-0190-5

[7] Dammak, S., Cecchini, M. J., Breadner, D., & Ward, A. D. (2023). Using deep learning to predict tumor mutational burden from multicenter H&E whole-slide images of lung squamous cell carcinoma. Journal of Medical Imaging, 10(1), 017502. https://doi.org/10.1117/1.JMI.10.1.017502

[8] Chen, S., Xiang, J., Wang, X., et al. (2022). Deep learningbased approach to reveal tumor mutational burden status from whole slide images across multiple cancer types. arXiv preprint arXiv:2204.03257. (No DOI)

[9] Lu, M. Y., Williamson, D. F. K., Chen, T. Y., & Mahmood, F. (2021). Data-efficient and weakly supervised computational pathology on whole-slide images. Nature Biomedical Engineering, 5(6), 555–570. https://doi.org/10.1038/s41551-020- 00682-w

[10] Shao, Z., Bian, H., Chen, Y., et al. (2021). TransMIL: Transformer-based multiple instance learning for whole slide image classification. In Advances in Neural Information Processing Systems (NeurIPS 34), 2136–2148. (No DOI) arXiv:2106.00908

[11] TRIPOD-AI Collaboration. (2024). Transparent reporting of a multivariable prediction model for AI/ML-based diagnosis and prognosis (TRIPOD-AI). BMJ, 386, e078378. https://doi. org/10.1136/bmj-2023-078378 Dean&Francis ISSN 2959-409X

[12] SPIRIT-AI Steering Group. (2020). SPIRIT-AI extension: Guidelines for clinical trial protocols involving artificial intelligence interventions. BMJ, 370, m3210. https://doi. org/10.1136/bmj.m3210

[13] CONSORT-AI Steering Group. (2020). CONSORT- AI extension: Reporting of clinical trials involving artificial intelligence interventions. BMJ, 370, m3164. https://doi. org/10.1136/bmj.m3164

[14] STARD-AI Steering Group. (2023). Reporting diagnostic accuracy studies that use AI: The STARD-AI protocol. BMJ Open, 13, e047709. https://doi.org/10.1136/ bmjopen-2023-047709

[15] Vickers, A. J., & Elkin, E. B. (2006). Decision curve analysis: A novel method for evaluating prediction models. Medical Decision Making, 26(6), 565–574. https://doi. org/10.1177/0272989X06295361

[16] DeLong, E. R., DeLong, D. M., & Clarke-Pearson, D. L. (1988). Comparing the areas under two or more correlated ROC curves: A nonparametric approach. Biometrics, 44(3), 837–845. https://doi.org/10.2307/2531595

[17] Harrell, F. E. (2015). Regression Modeling Strategies (2nd ed.). Springer. https://doi.org/10.1007/978-3-319-19425-7

[18] Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (ICML 2017), PMLR 70, 1321–1330. (No DOI) http://proceedings.mlr. press/v70/guo17a.html

[19] DECIDE-AI Steering Group. (2022). DECIDE-AI: Reporting guidelines for early-stage clinical evaluation of AI- based decision support systems. Nature Medicine, 28, 924–933. https://doi.org/10.1038/s41591-022-01772-9

[20] Lu, W., Toss, M., Dawood, M., Rakha, E., Rajpoot, N., & Minhas, F. (2022). SlideGraph+: Whole-slide image-level graphs to predict HER2 status in breast cancer. Medical Image Analysis, 80, 102486. https://doi.org/10.1016/j.media.2022.102486

[21] Chen, C., Lu, M. Y., Williamson, D. F. K., et al. (2022). Self-supervised instance-level retrieval for computational pathology. Nature Biomedical Engineering, 6, 1420–1434. https://doi.org/10.1038/s41551-022-00929-8

[22] Xu, H., Usuyama, N., Bagga, J., et al. (2024). A wholeslide foundation model for digital pathology from real-world data. Nature, 629, 358–365. https://doi.org/10.1038/s41586-024- 07441-w

[23] Wang, X., Yang, S., Zhang, J., et al. (2024). A pathology foundation model for cancer diagnosis and prognosis. Nature, 630, 131–138. https://doi.org/10.1038/s41586-024-07894-z

[24] Chen, R.-J., Ding, T., Williamson, D. F. K., et al. (2024). A generalist pathology foundation model. Nature Medicine, 30, 1703–1713. https://doi.org/10.1038/s41591-024-02857-3

Downloads

Published

2025-10-23