Rxivist logo

Widespread false gene gains caused by duplication errors in genome assemblies

By Byung June Ko, Chul Lee, Juwan Kim, Arang Rhie, DongAhn Yoo, Kerstin Howe, Jonathan Wood, Seoae Cho, Samara Brown, Giulio Formenti, Erich D Jarvis, Heebal Kim

Posted 09 Apr 2021
bioRxiv DOI: 10.1101/2021.04.09.438957

False duplications in genome assemblies lead to false biological conclusions. We quantified false duplications in previous genome assemblies and their new counterparts of the same species (platypus, zebra finch, Anna's hummingbird) generated by the Vertebrate Genomes Project (VGP). Whole genome alignments revealed that 4 to 16% of the sequences were falsely duplicated in the previous assemblies, impacting hundreds to thousands of genes. These led to overestimated gene family expansions. The main source of the false duplications was heterotype duplications, where the haplotype sequences were more divergent than other parts of the genome leading the assembly algorithms to classify them as separate genes or genomic regions. A minor source was sequencing errors. Although present in a smaller proportion, we observed false duplications remaining in the VGP assemblies that can be identified and purged. This study highlights the need for more advanced assembly methods that better separates haplotypes and sequence errors, and the need for cautious analyses on gene gains.

Download data

  • Downloaded 640 times
  • Download rankings, all-time:
    • Site-wide: 49,282
    • In genomics: 3,809
  • Year to date:
    • Site-wide: 8,488
  • Since beginning of last month:
    • Site-wide: 38,671

Altmetric data

Downloads over time

Distribution of downloads per paper, site-wide