Estimating time of HIV-1 infection from next-generation sequence diversity

Vadim Puller; Richard Neher; Jan Albert

doi:10.1371/journal.pcbi.1005775

Estimating time of HIV-1 infection from next-generation sequence diversity

PLoS Comput Biol. 2017 Oct 2;13(10):e1005775. doi: 10.1371/journal.pcbi.1005775. eCollection 2017 Oct.

Authors

Vadim Puller^{1

2

3}, Richard Neher^{1

2

3}, Jan Albert^{4

5}

Affiliations

¹ Max Planck Institute for Developmental Biology, Tübingen, Germany.
² Biozentrum, University of Basel, Basel, Switzerland.
³ SIB Swiss Institute of Bioinformatics, Basel, Switzerland.
⁴ Department of Microbiology, Tumor and Cell Biology, Karolinska Institute, Stockholm, Sweden.
⁵ Department of Clinical Microbiology, Karolinska University Hospital, Stockholm, Sweden.

Abstract

Estimating the time since infection (TI) in newly diagnosed HIV-1 patients is challenging, but important to understand the epidemiology of the infection. Here we explore the utility of virus diversity estimated by next-generation sequencing (NGS) as novel biomarker by using a recent genome-wide longitudinal dataset obtained from 11 untreated HIV-1-infected patients with known dates of infection. The results were validated on a second dataset from 31 patients. Virus diversity increased linearly with time, particularly at 3rd codon positions, with little inter-patient variation. The precision of the TI estimate improved with increasing sequencing depth, showing that diversity in NGS data yields superior estimates to the number of ambiguous sites in Sanger sequences, which is one of the alternative biomarkers. The full advantage of deep NGS was utilized with continuous diversity measures such as average pairwise distance or site entropy, rather than the fraction of polymorphic sites. The precision depended on the genomic region and codon position and was highest when 3rd codon positions in the entire pol gene were used. For these data, TI estimates had a mean absolute error of around 1 year. The error increased only slightly from around 0.6 years at a TI of 6 months to around 1.1 years at 6 years. Our results show that virus diversity determined by NGS can be used to estimate time since HIV-1 infection many years after the infection, in contrast to most alternative biomarkers. We provide the regression coefficients as well as web tool for TI estimation.

MeSH terms

Genome, Viral / genetics*
Genomics / methods*
HIV Infections / epidemiology*
HIV Infections / virology*
HIV-1 / genetics*
High-Throughput Nucleotide Sequencing
Humans
Models, Statistical
Sequence Analysis, DNA / methods*

Grants and funding

This work was supported by: European Research Council (https://erc.europa.eu/), Stg. 260686, principal investigator RN; Swedish Research Council (https://www.vr.se/), K2014-57X-09935, principal investigator JA. The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.