Rxivist logo

Rxivist combines preprints from bioRxiv with data from Twitter to help you find the papers being discussed in your field. Currently indexing 70,897 bioRxiv papers from 309,427 authors.

Cooler: scalable storage for Hi-C data and other genomically-labeled arrays

By Nezar Abdennur, Leonid Mirny

Posted 22 Feb 2019
bioRxiv DOI: 10.1101/557660 (published DOI: 10.1093/bioinformatics/btz540)

Most existing coverage-based (epi)genomic datasets are one-dimensional, but newer technologies probing interactions (physical, genetic, etc.) produce quantitative maps with two-dimensional genomic coordinate systems. Storage and computational costs mount sharply with data resolution when such maps are stored in dense form. Hence, there is a pressing need to develop data storage strategies that handle the full range of useful resolutions in multidimensional genomic datasets by taking advantage of their sparse nature, while supporting efficient compression and providing fast random access to facilitate development of scalable algorithms for data analysis. We developed a file format called cooler, based on a sparse data model, that can support genomically-labeled matrices at any resolution. It has the flexibility to accommodate various descriptions of the data axes (genomic coordinates, tracks and bin annotations), resolutions, data density patterns, and metadata. Cooler is based on HDF5 and is supported by a Python library and command line suite to create, read, inspect and manipulate cooler data collections. The format has been adopted as a standard by the NIH 4D Nucleome Consortium. Cooler is cross-platform, BSD-licensed, and can be installed from the Python Package Index or the bioconda repository. The source code is maintained on Github at https://github.com/mirnylab/cooler.

Download data

  • Downloaded 968 times
  • Download rankings, all-time:
    • Site-wide: 8,768 out of 70,897
    • In bioinformatics: 1,517 out of 6,940
  • Year to date:
    • Site-wide: 4,774 out of 70,897
  • Since beginning of last month:
    • Site-wide: 8,747 out of 70,897

Altmetric data


Downloads over time

Distribution of downloads per paper, site-wide


PanLingua

Sign up for the Rxivist weekly newsletter! (Click here for more details.)


News