Project ArHa: Arenaviruses & Hantaviruses

A global synthesis of rodent-borne viral pathogens

Project ArHa brings together published records of small mammals sampled for arenaviruses and hantaviruses, worldwide. I lead it with Stephanie Seifert, with support from the Fellows-in-Residence programme of the Verena consortium, and in collaboration with David Redding and his group at the Natural History Museum in London.

Each record gives the species sampled, the place and date, the assays used and their results. Records where no animal tested positive are retained in the database, so it describes sampling effort as well as detection. That is what makes the biases in surveillance measurable.

54,865 host sampling records in the v1.1 release

103 countries with sampling records

46% of rodent genera never sampled for these viruses

Use the data

Where hosts have been sampled

Host sampling records per hexagon

    Loading the sampling record.

    A record is one host species sampled at one place and time. Records where no animals of that species were caught still count, so a hexagon can hold more records than individuals.

    Drag to rotate, or focus the globe and use the arrow keys. The records themselves are in the ArHa database explorer.

    Show as a table
    Host sampling records by continent and country.
    Continent Country Records Locations Individuals

    Host status and surveillance bias

    Host status for arenaviruses and hantaviruses appears to track synanthropy and a fast pace of life. Surveillance is not random, so these associations could be artefacts of where and what has been sampled. The ArHa preprint uses the database’s record of sampling effort to assess whether host status is predictable once that effort is accounted for.

    Surveillance intensity follows night-time light intensity and accessibility, not local host richness. Forty-six per cent of rodent genera have never been sampled.

    World map of surveillance residuals: districts sampled more or less intensely than predicted by night-time light, accessibility, population density and host richness.

    Global surveillance inequality. Residuals from a zero-inflated negative binomial model of surveillance effort. Orange districts are sampled more than night-time light, accessibility, population density and local host richness predict; blue districts less. Figure 1 of the preprint, CC BY 4.0.

    Bayesian phylogenetic mixed models of 43,677 host–virus pairs estimate host status after adjusting for sampling effort. Faster-lived species tended to be detected as hosts more often, though the credible interval included zero. Synanthropy was independently associated with host status, with about twice the odds for obligate commensals. Withheld continents were predicted with an AUC of 0.81 to 0.88.

    Four panels: posterior log-odds for synanthropy, pace of life and sampling effort; probability of host status rising with individuals sampled; falling slightly from fast to slow pace of life; and rising with synanthropy.

    Associations with host status. Posterior log-odds for sampling effort, pace of life and synanthropy, and the marginal effect of each on the probability of host status. Figure 2 of the preprint, CC BY 4.0.

    Projected globally, community host probability varies largely independently of species richness (R² = 0.028). Diverse assemblages do not carry systematically higher host probability.

    Bivariate world map of small mammal species richness against community host probability at 20 km resolution.

    Community host probability. Species richness against community host probability, the mean predicted probability of host status across the small mammals present in each 20 km cell. Figure 3 of the preprint, CC BY 4.0.

    Host and virus phylogenies are congruent in both families, but most associations depart from strict co-divergence. Host switching has occurred alongside shared evolutionary history.

    Tanglegram linking the mammalian host phylogeny above to hantavirus and arenavirus phylogenies below, with many crossing links.

    Phylogenetic congruence between viruses and their hosts. The mammalian host phylogeny (top) against the Hantaviridae (bottom left) and Arenaviridae (bottom right), linked to PCR-positive host species. Thicker, darker links mark stronger congruence; thinner, fainter links suggest host switching. Figure 4 of the preprint, CC BY 4.0.

    Papers

    Molecular and ecological determinants of effective reassortment in orthohantaviruses
    Rivero R, Simons D, Damodaran L, Karegi I, Gurev S, Becker DJ, et al.
    Preprint, 2026 · PDF (preprint)

    Summary

    Segmented RNA viruses can exchange whole genome segments when two lineages infect the same host, but most reassortants never establish. Why some persist and others do not is an open problem in viral evolution, and orthohantaviruses offer a tractable case because host associations are well described.

    Reassortant histories were reconstructed across 553 genomes from seven orthohantavirus species sampled between 1983 and 2024, using phylogenetic reconciliation and molecular dating, then modelled with Bayesian hierarchical models. Retained reassortment varied by species, absent in Andes virus and frequent in Dobrava-Belgrade, Sin Nombre, Seoul, Puumala and Tula viruses, so it is not a genus-wide constant. Local host overlap was the strongest ecological correlate, while cross-segment linkage and terminal RNA structure acted as a molecular filter.

    Establishment was most probable where ecological opportunity coincided with molecular permissiveness, their interaction being the strongest signal in the establishment models (posterior probability 0.97). Reassortment therefore appears to be sequentially filtered: lineages must first meet in a host, then exchange compatible segments, then land in a lineage background that permits establishment.

    Viral reservoir status in small mammals emerges as a predictable life-history trait after correcting for surveillance bias
    Simons D, Rivero R, Rickard G, Martinez-Checa A, Gordon H, Redding DW, Seifert SN
    Preprint, 2026

    Summary

    Four panels showing posterior log-odds for synanthropy, pace of life and sampling effort, and the marginal effect of each on the probability of host status.

    Associations with host status after adjusting for sampling effort.

    Small mammals are the principal hosts of arenaviruses and hantaviruses, but global surveillance is non-random, so the determinants of host status remain contested. Associations with synanthropy or life history could reflect where and which species have been sampled rather than host biology.

    Using the ArHa database, 729 studies and 695,000 diagnostic assays across 637 species were harmonised. Surveillance intensity followed night-time light intensity and accessibility rather than local host richness, and 46% of rodent genera have never been sampled. Bayesian phylogenetic mixed models of 43,677 host–virus pairs, adjusted for sampling effort, estimate that faster-lived species tended to be detected as hosts more often, though the credible interval included zero (pd 93.2%). Synanthropy was independently associated with host status, with obligate commensals at about twice the odds. Predictions discriminated hosts in withheld continents (AUC 0.81 to 0.88).

    Projected globally, community host probability varied largely independently of species richness (R² = 0.028). Host and virus phylogenies were congruent, with host switching alongside co-divergence.

    Protocol to produce a systematic Arenavirus and Hantavirus host-pathogen database: Project ArHa.
    Simons D, Rivero R, Guiote AMC, Gordon HLM, Milne GC, Rickard G, Redding DW, Seifert SN
    Preprint, 2025 · Wellcome Open Research, 2025 · PDF

    Summary

    World map of small-mammal sampling locations for arenaviruses and hantaviruses, points coloured by virus family, with dense coverage in Europe and North America and sparse coverage across Africa and Asia.

    Sampling locations by virus family. The holes in the coverage are the finding.

    Arenaviruses and hantaviruses are hosted primarily by rodents and shrews and cause substantial human morbidity, yet their global distribution is known only through a literature scattered across decades, languages and disciplines. Without a synthesis, sampling effort cannot be separated from pathogen occurrence.

    This protocol specifies how Project ArHa assembles that synthesis: the search strategy across bibliographic databases, screening and inclusion criteria, extraction fields, and a relational structure of five tables covering citations, study design, host occurrence, pathogen assay and genetic sequence. Critically, non-detections are retained, so absence of virus in a sampled host is recorded rather than lost.

    The resulting database is spatially and temporally explicit, which makes it usable for quantifying surveillance bias as well as for modelling host-pathogen association. An accompanying Shiny application opens the data to users who do not code. The protocol was published before extraction completed, fixing the method in advance.

    Last updated 6 October 2026