The Impact of Digital Histopathology Batch Effect on Deep Learning Model Accuracy and Bias
Frederick Matthew Howard,
Olufunmilayo I. Olopade,
Alexander T Pearson
Posted 04 Dec 2020
bioRxiv DOI: 10.1101/2020.12.03.410845
Posted 04 Dec 2020
The Cancer Genome Atlas (TCGA) is one of the largest biorepositories of digital histology. Deep learning (DL) models have been trained on TCGA to predict numerous features directly from histology, including survival, gene expression patterns, and driver mutations. However, we demonstrate that these features vary substantially across tissue submitting sites in TCGA for over 3,000 patients with six cancer subtypes. We demonstrate that staining differences between submitting sites exist and can easily be identified by DL models. We quantify the digital image characteristics constituting this histologic batch effect. We demonstrate that DL models can accurately predict submitting site with an area under the receiver operating characteristic curve of 0.850 or higher despite commonly used stain normalization and augmentation methods. To correct for biased estimates of accuracy, we propose a mixed integer quadratic programming method to separate data into cross folds to abrogate this bias. As a demonstration, we show that patient ethnicity can be inferred from histology due to site-level staining patterns; this must be accounted for to ensure equitable applicability of deep learning models.
- Downloaded 344 times
- Download rankings, all-time:
- Site-wide: 76,766
- In bioinformatics: 7,103
- Year to date:
- Site-wide: 15,038
- Since beginning of last month:
- Site-wide: 11,410
Downloads over time
Distribution of downloads per paper, site-wide
- 27 Nov 2020: The website and API now include results pulled from medRxiv as well as bioRxiv.
- 18 Dec 2019: We're pleased to announce PanLingua, a new tool that enables you to search for machine-translated bioRxiv preprints using more than 100 different languages.
- 21 May 2019: PLOS Biology has published a community page about Rxivist.org and its design.
- 10 May 2019: The paper analyzing the Rxivist dataset has been published at eLife.
- 1 Mar 2019: We now have summary statistics about bioRxiv downloads and submissions.
- 8 Feb 2019: Data from Altmetric is now available on the Rxivist details page for every preprint. Look for the "donut" under the download metrics.
- 30 Jan 2019: preLights has featured the Rxivist preprint and written about our findings.
- 22 Jan 2019: Nature just published an article about Rxivist and our data.
- 13 Jan 2019: The Rxivist preprint is live!