Pangenome: a novel frontier for genetics
Figure 1. The pangenome construction, from the selection of 47 individuals, to the sequence of 94 haploid genomes and the assembled graph. The use of pangenome as a reference in a mapping workflow to discovery genetic variations.
Abstract
The human pangenome, a new comprehensive representation of human genetic diversity, offers a fresh perspective for genetic studies while paving the way for personalized medicine. Could the pangenome become the next genome reference by eliminating the current reference bias?
Review
Introduction
In 2023, a new chapter in human genomics began with the release of the first draft human pangenome by the Human Pangenome Reference Consortium (HPRC) [1]. This pangenome represents the combined genetic diversity of 47 individuals from across the globe, offering a more comprehensive picture compared to single-reference genomes like GRCh38, released in 2003, and CHM13, released in 2022. While these single references have been invaluable, they inherently struggle to capture the full spectrum of human genetic variation. A single genome reference cannot represent the genomic landscape of the diverse human population. The main issue with using a single reference genome is the reference bias, which significantly impacts variant discovery, gene-disease association studies, and the accuracy of genetic analyses.
By creating a pangenome, the authors gather diverse genomes from various individuals and collapsed identical segments into a single copy, while preserving unique segments as distinct regions of the DNA. This process allows to capture the collective genetic variation, providing a comprehensive representation of both shared and individual-specific genomic sequences.
Pangenome construction
Liao and colleagues have faced two major challenges by constructing the human pangenome: obtaining very high-quality assemblies and find a way to represent these genomes.
The first step to represent human population is the selection of representative samples. The pangenome contains 47 phased diploid assemblies, with 18 of them selected from previous studies and 29 samples selected from the 1000 Genomes project (1KG). To create diploid phased assemblies, parent-child trios were identified. Subsequently, the most representative individual from each subpopulation was defined by conducting a principal component analysis (PCA) on each chromosome. Then they used long and short reads sequencing methods to achieve high-quality and nearly complete assemblies of the sample genomes. However, a more recent study [2] reported that the assembly of diploid genomes necessitates specific software (Hifiasm) that struggles with repetitive regions, leading to gaps in the assembly. Such variability and incompletely assembled regions are important targets for future algorithmic development and pangenome representation.
The process of generating a combined pangenome is the current target of the research and constitutes the second challenge in pangenome construction. Liao and colleagues employed three distinct graph algorithmic construction methods: Minigraph, Minigraph-Cactus (MC), and Pangenome Graph Builder (PGGB). The primary distinction among these approaches is that Minigraph and MC incorporate GRCh38 as a reference assembly and progressively integrate additional variants, whereas PGGB is reference-free. The different techniques influence the properties of the pangenome graph. PGGB creates a larger and more complex map that includes every variation, whereas the other two methods are more linear and easier to use but smaller, representing fewer variations.

Figure 2. H1 and H2 are two haplotypes that differ in the copy number of a chromosomal segment. The three segments (S1, S2, S3) present small variations such as SNPs and indels. While Minigraph includes only structural variations of at least 50 bp, MC adds smaller variations (less than 50 bp). Finally, PGGB collapses the segmental variation into loops, including all variations.
The HPRC have publicly released 94 de novo assemblies and the three pangenome graphs available on their website. The current usability of these graphs is constrained by the complexity of the interface and of the models. The software remains difficult to approach and is still in a preliminary stage.
Variant discovery
The utilization of pangenome graphs as a reference for detecting small and structural variants leads to an increased detection rate compared to the single-genome reference GRCH38. The detection rate of structural variants (SV) increases up to 104%, while it reduces the error rate of individuate small variants by 34%. This approach enhances the sensitivity and accuracy of variant identification across various genomic alterations.
The detection of SVs has been an issue due to the lack of alternative alleles in traditional reference genomes, despite their critical impact on gene function. This limitation results in missing over two-thirds of SVs in studies utilizing short-read data and standard reference assemblies. To address this reference bias, the pangenome approach was employed, leading to the discovery of numerous SVs, especially within complex loci. These complex loci, characterized by multiple structural alleles, encompass regions of significant medical relevance. The study identified 44 complex SV sites that overlapped with protein-coding genes of medical importance, such as RHD and RHCE, where novel haplotypes, including duplication and inversion alleles, were uncovered.
Discussion
The pangenome comprises core genes (shared between all individual of a species), and population- or individual-specific genes. Core genes, are essential for basic life processes, influence the biological functions and phenotypic traits of a species. Population- or individual-specific genes, on the other hand, can be involved in modulating secondary metabolism or adapting to environmental challenges [3]. Therefore, a near-term application of the pangenome will be in research to identify previously missing non-shared genetic components, such as large structural variants (SVs) and presence-absence variations (PAVs). Integrating a pangenome reference into read mapping workflows in research could be easier with respect to its usage in medical practice.
The pangenome reference promises to be a solution to overcome reference bias, leading to a more comprehensive understanding of human genetics and a significant advancement in research and making it more representative of individuals from diverse ancestries. Nonetheless, it is crucial to acknowledge that the specific composition of the sample set could potentially influence these outcomes. Consequently, an analysis involving a larger and more diverse sample size is essential for drawing robust conclusions.
In the future the pangenome could become a valuable resource for personalized medicine, revolutionizing the way how diseases are diagnosed and prevented. By incorporating the diverse genetic information captured in the pangenome, healthcare professionals can tailor treatments and interventions to individual genetic profiles, leading to more precise and effective healthcare strategies. In this context, one limitation of the current pangenome is the absence of sufficient clinical data regarding the individuals incorporated in the reference. Establishing the medical implications for each individual could enhance the comprehensiveness of the data.
Conclusions
The pangenome graphs demonstrate a high concordance with standard reference genomes like GRCH38 and CHm13, while also exhibiting superior performance in variant discovery. The pangenome emerges as a valuable tool in genomics research, poised to revolutionize our understanding, analysis, and prevention of genetic diseases, as well as human genome diversity.
Despite its potential, the current pangenome still requires an increase in the number of samples and enhancement of assembly quality, particularly in challenging and repetitive genomic regions. Future advancements in sequencing and assembly technologies hold promise for overcoming these limitations. The Human Pangenome Reference Consortium (HPRC) aims to address these challenges by expanding the pangenome to include data from 350 individuals and enhancing sequencing quality up to the level of the T2T-CHM13 assemble [4].
Nevertheless, the current lack of user-friendly tools presents a challenge to its widespread utilization across various fields. Integration and practical application of the pangenome will thus require several years. The value of this project lies in its potential to establish new standards for capturing variant diversity, creating a highly representative and complete reference of the human genome, reflecting the global genetic diversity of humanity.
References
- Liao, WW., Asri, M., Ebler, J. et al. A draft human pangenome reference. Nature May;617(7960):312-324.
- Porubsky D, Vollger MR, Harvey WT, et al. Gaps and complex structurally variant loci in phased genome assemblies. Genome Res. 2023. 2023 Apr;33(4):496-510.
- Yu, Y., Chen, H. Human pangenome: far-reaching implications in precision medicine. Front. Med. 2023 Dec. 29
- Nurk S, Koren S, Rhie., et al. The complete sequence of a human genome. Science. 2022 Apr;376(6588):44-53.
