Machine learning reduced gene/non-coding RNA features that classify Schizophrenia patients accurately and highlight insightful gene clusters
By
Yichuan Liu,
Huiqi Qu,
Xiao Chang,
Lifeng Tian,
Joseph Glessner,
Patrick A Sleiman,
Hakon Hakonarson
Posted 12 Jun 2020
medRxiv DOI: 10.1101/2020.06.08.20125906
Schizophrenia (SCZ) is a chronic and severely disabling neurodevelopmental disorder that affects people worldwide. RNA-seq has been a powerful method to detect the differentially expressed genes/non-coding RNAs in patients; however, due to overfitting problems differentially expressed targets (DETs) cannot be used properly as biomarkers. In this study, dorsolateral prefrontal cortex (dlpfc) RNA-seq data from 254 individuals was obtained from the CommonMind consortium and analyzed with machine learning methods, including random forest, forward feature selection (ffs), and factor analysis, to reduce the numbers of gene/non-coding RNA feature vectors to overcome overfitting problem and explore involved functional clusters. In 2-fold shuffle testing, the average predictive accuracy for SCZ patients was 67% based on coding genes, and the 96% based on long non-coding RNAs (lncRNAs). Coding genes were further clustered into 14 factors and lncRNAs were clustered into 45 factors to represent the underlying features. The largest contribution factor for coding genes contains number of genes critical in neurodevelopment and previously reported in relation with various brain disorders. Genomic loci of lncRNAs were more insightful, enriched for genes critical in synapse function (p=7.3E-3), cell junction (p=0.017), neuron differentiation (p=8.3E-3), phosphorylation (8.2E-4), and involving the Wnt signaling pathway (p=0.029). Taken together, machine learning is a powerful algorithm to reduce functional biomarkers in SCZ patients. The lncRNAs capture the characteristics of SCZ tissue more accurately than mRNA as the formers regulate every level of gene expression, not limited to mRNA levels.
Download data
- Downloaded 185 times
- Download rankings, all-time:
- Site-wide: 110,074
- In psychiatry and clinical psychology: 451
- Year to date:
- Site-wide: 72,165
- Since beginning of last month:
- Site-wide: 63,620
Altmetric data
Downloads over time
Distribution of downloads per paper, site-wide
PanLingua
News
- 27 Nov 2020: The website and API now include results pulled from medRxiv as well as bioRxiv.
- 18 Dec 2019: We're pleased to announce PanLingua, a new tool that enables you to search for machine-translated bioRxiv preprints using more than 100 different languages.
- 21 May 2019: PLOS Biology has published a community page about Rxivist.org and its design.
- 10 May 2019: The paper analyzing the Rxivist dataset has been published at eLife.
- 1 Mar 2019: We now have summary statistics about bioRxiv downloads and submissions.
- 8 Feb 2019: Data from Altmetric is now available on the Rxivist details page for every preprint. Look for the "donut" under the download metrics.
- 30 Jan 2019: preLights has featured the Rxivist preprint and written about our findings.
- 22 Jan 2019: Nature just published an article about Rxivist and our data.
- 13 Jan 2019: The Rxivist preprint is live!