VGsim : scalable viral genealogy simulator for global pandemic
Accurate simulation of complex biological processes is an essential component of developing and validating new technologies and inference approaches. As an effort to help contain the COVID-19 pandemic, large numbers of SARS-CoV-2 genomes have been sequenced from most regions in the world. More than 5.5 million viral sequences are publicly available as of November 2021. Many studies estimate viral genealogies from these sequences, as these can provide valuable information about the spread of the pandemic across time and space. Additionally such data are a rich source of information about molecular evolutionary processes including natural selection, for example allowing the identification of new variants with transmissibility and immunity evasion advantages. To our knowledge, there is no framework that is both efficient and flexible enough to simulate the pandemic to approximate world-scale scenarios and generate viral genealogies of millions of samples. Here, we introduce a new fast simulator VGsim which addresses the problem of simulation genealogies under epidemiological models. The simulation process is split into two phases. During the forward run the algorithm generates a chain of population-level events reflecting the dynamics of the pandemic using an hierarchical version of the Gillespie algorithm. During the backward run a coalescent-like approach generates a tree genealogy of samples conditioning on the population-level events chain generated during the forward run. Our software can model complex population structure, epistasis and immunity escape. The code is freely available at https://github.com/Genomics-HSE/VGsim.
Errataetall: |
UpdateIn: PLoS Comput Biol. 2022 Aug 24;18(8):e1010409. - PMID 36001646 |
---|---|
Medienart: |
E-Artikel |
Erscheinungsjahr: |
2021 |
---|---|
Erschienen: |
2021 |
Enthalten in: |
Zur Gesamtaufnahme - year:2021 |
---|---|
Enthalten in: |
medRxiv : the preprint server for health sciences - (2021) vom: 02. Dez. |
Sprache: |
Englisch |
---|
Beteiligte Personen: |
Shchur, Vladimir [VerfasserIn] |
---|
Links: |
---|
Themen: |
---|
Anmerkungen: |
Date Revised 08.11.2023 published: Electronic UpdateIn: PLoS Comput Biol. 2022 Aug 24;18(8):e1010409. - PMID 36001646 Citation Status PubMed-not-MEDLINE |
---|
doi: |
10.1101/2021.04.21.21255891 |
---|
funding: |
|
---|---|
Förderinstitution / Projekttitel: |
|
PPN (Katalog-ID): |
NLM32501549X |
---|
LEADER | 01000naa a22002652 4500 | ||
---|---|---|---|
001 | NLM32501549X | ||
003 | DE-627 | ||
005 | 20231225191534.0 | ||
007 | cr uuu---uuuuu | ||
008 | 231225s2021 xx |||||o 00| ||eng c | ||
024 | 7 | |a 10.1101/2021.04.21.21255891 |2 doi | |
028 | 5 | 2 | |a pubmed24n1083.xml |
035 | |a (DE-627)NLM32501549X | ||
035 | |a (NLM)33948608 | ||
035 | |a (PII)2021.04.21.21255891 | ||
040 | |a DE-627 |b ger |c DE-627 |e rakwb | ||
041 | |a eng | ||
100 | 1 | |a Shchur, Vladimir |e verfasserin |4 aut | |
245 | 1 | 0 | |a VGsim |b scalable viral genealogy simulator for global pandemic |
264 | 1 | |c 2021 | |
336 | |a Text |b txt |2 rdacontent | ||
337 | |a ƒaComputermedien |b c |2 rdamedia | ||
338 | |a ƒa Online-Ressource |b cr |2 rdacarrier | ||
500 | |a Date Revised 08.11.2023 | ||
500 | |a published: Electronic | ||
500 | |a UpdateIn: PLoS Comput Biol. 2022 Aug 24;18(8):e1010409. - PMID 36001646 | ||
500 | |a Citation Status PubMed-not-MEDLINE | ||
520 | |a Accurate simulation of complex biological processes is an essential component of developing and validating new technologies and inference approaches. As an effort to help contain the COVID-19 pandemic, large numbers of SARS-CoV-2 genomes have been sequenced from most regions in the world. More than 5.5 million viral sequences are publicly available as of November 2021. Many studies estimate viral genealogies from these sequences, as these can provide valuable information about the spread of the pandemic across time and space. Additionally such data are a rich source of information about molecular evolutionary processes including natural selection, for example allowing the identification of new variants with transmissibility and immunity evasion advantages. To our knowledge, there is no framework that is both efficient and flexible enough to simulate the pandemic to approximate world-scale scenarios and generate viral genealogies of millions of samples. Here, we introduce a new fast simulator VGsim which addresses the problem of simulation genealogies under epidemiological models. The simulation process is split into two phases. During the forward run the algorithm generates a chain of population-level events reflecting the dynamics of the pandemic using an hierarchical version of the Gillespie algorithm. During the backward run a coalescent-like approach generates a tree genealogy of samples conditioning on the population-level events chain generated during the forward run. Our software can model complex population structure, epistasis and immunity escape. The code is freely available at https://github.com/Genomics-HSE/VGsim | ||
650 | 4 | |a Preprint | |
700 | 1 | |a Spirin, Vadim |e verfasserin |4 aut | |
700 | 1 | |a Sirotkin, Dmitry |e verfasserin |4 aut | |
700 | 1 | |a Burovski, Evgeni |e verfasserin |4 aut | |
700 | 1 | |a De Maio, Nicola |e verfasserin |4 aut | |
700 | 1 | |a Corbett-Detig, Russell |e verfasserin |4 aut | |
773 | 0 | 8 | |i Enthalten in |t medRxiv : the preprint server for health sciences |d 2020 |g (2021) vom: 02. Dez. |w (DE-627)NLM310900166 |7 nnns |
773 | 1 | 8 | |g year:2021 |g day:02 |g month:12 |
856 | 4 | 0 | |u http://dx.doi.org/10.1101/2021.04.21.21255891 |3 Volltext |
912 | |a GBV_USEFLAG_A | ||
912 | |a GBV_NLM | ||
951 | |a AR | ||
952 | |j 2021 |b 02 |c 12 |