Rxivist logo

Predicting the spectrum of TCR repertoire sharing with a data-driven model of recombination

By Yuval Elhanati, Zachary Sethna, Curtis G Callan, Thierry Mora, Aleksandra M Walczak

Posted 02 Mar 2018
bioRxiv DOI: 10.1101/275602 (published DOI: 10.1111/imr.12665)

Despite the extreme diversity of T cell repertoires, many identical T-cell receptor (TCR) sequences are found in a large number of individual mice and humans. These widely-shared sequences, often referred to as 'public', have been suggested to be over-represented due to their potential immune functionality or their ease of generation by V(D)J recombination. Here we show that even for large cohorts the observed degree of sharing of TCR sequences between individuals is well predicted by a model accounting for by the known quantitative statistical biases in the generation process, together with a simple model of thymic selection. Whether a sequence is shared by many individuals is predicted to depend on the number of queried individuals and the sampling depth, as well as on the sequence itself, in agreement with the data. We introduce the degree of publicness conditional on the queried cohort size and the size of the sampled repertoires. Based on these observations we propose a public/private sequence classifier, 'PUBLIC' (Public Universal Binary Likelihood Inference Classifier), based on the generation probability, which performs very well even for small cohort sizes.

Download data

  • Downloaded 660 times
  • Download rankings, all-time:
    • Site-wide: 23,192 out of 94,912
    • In immunology: 613 out of 2,857
  • Year to date:
    • Site-wide: 28,340 out of 94,912
  • Since beginning of last month:
    • Site-wide: 28,616 out of 94,912

Altmetric data

Downloads over time

Distribution of downloads per paper, site-wide


Sign up for the Rxivist weekly newsletter! (Click here for more details.)