Benegas, G., Ye, C., Albors, C., Li, J. C. & Song, Y. S. Genomic language models: opportunities and challenges. Trends Genet. 41, 286–302 (2025). Article CAS PubMed Google Scholar Benegas, G., Eraslan, G. & Song, Y. S. Benchmarking DNA sequence models for causal regulatory variant prediction in human genetics. Preprint at bioRxiv https://doi.org/10.1101/2025.02.11.637758 (2025). Dalla-Torre,
Benegas, G., Ye, C., Albors, C., Li, J. C. & Song, Y. S. Genomic language models: opportunities and challenges. Trends Genet. 41, 286–302 (2025).
Google Scholar
Benegas, G., Eraslan, G. & Song, Y. S. Benchmarking DNA sequence models for causal regulatory variant prediction in human genetics. Preprint at bioRxiv https://doi.org/10.1101/2025.02.11.637758 (2025).
Dalla-Torre, H. et al. Nucleotide transformer: building and evaluating robust foundation models for human genomics. Nat. Methods 22, 287–297 (2025).
Google Scholar
Brixi, G. et al. Genome modelling and design across all domains of life with Evo 2. Nature 652, 1349–1361 (2026).
Google Scholar
Clarke, B. et al. Integration of variant annotations using deep set networks boosts rare variant association testing. Nat. Genet. 56, 2271–2280 (2024).
Google Scholar
Dayhoff, M. O., Schwartz, R. M. & Orcutt, B. C. A model of evolutionary change in proteins. Atlas Protein Seq. Struct. 5, 345–352 (1978).
Sullivan, P. F. et al. Leveraging base-pair mammalian constraint to understand genetic variation and human disease. Science 380, eabn2937 (2023).
Google Scholar
Kuderna, L. F. et al. Identification of constrained sequence elements across 239 primate genomes. Nature 625, 735–742 (2024).
Google Scholar
Benegas, G., Batra, S. S. & Song, Y. S. DNA language models are powerful predictors of genome-wide variant effects. Proc. Natl Acad. Sci. USA 120, e2311219120 (2023).
Google Scholar
Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).
Google Scholar
Rao, R. M. et al. MSA Transformer. In Proc. 38th Int. Conf. on Machine Learning vol. 139 (eds Meila, M. & Zhang, T.) 8844–8856 (PMLR, 2021).
Frazer, J. et al. Disease variant prediction with deep generative models of evolutionary data. Nature 599, 91–95 (2021).
Google Scholar
Truong, T. Jr & Bepler, T. PoET: a generative model of protein families as sequences-of-sequences. Adv. Neural Info. Process. Syst. 36, 77379–77415 (2023).
Yang, K. K. et al. The Dayhoff Atlas: scaling sequence diversity for improved protein generation. Preprint at bioRxiv https://doi.org/10.1101/2025.07.21.665991 (2025).
Akiyama, Y. et al. Expanding the scope of protein language modeling to protein-protein interactions with MSA pairformer. Cell 189, 4964–4979 (2026).
Google Scholar
Blanchette, M. et al. Aligning multiple genomic sequences with the threaded blockset aligner. Genome Res. 14, 708–715 (2004).
Google Scholar
Armstrong, J. et al. Progressive Cactus is a multiple-genome aligner for the thousand-genome era. Nature 587, 246–251 (2020).
Google Scholar
Siepel, A. et al. Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes. Genome Res. 15, 1034–1050 (2005).
Google Scholar
Pollard, K. S., Hubisz, M. J., Rosenbloom, K. R. & Siepel, A. Detection of nonneutral substitution rates on mammalian phylogenies. Genome Res. 20, 110–121 (2010).
Google Scholar
Christmas, M. J. et al. Evolutionary constraint and innovation across hundreds of placental mammals. Science 380, eabn3943 (2023).
Google Scholar
Rhie, A. et al. Towards complete and error-free genome assemblies of all vertebrate species. Nature 592, 737–746 (2021).
Google Scholar
Benegas, G., Albors, C., Aw, A. J., Ye, C. & Song, Y. S. A DNA language model based on multispecies alignment predicts the effects of genome-wide variants. Nat. Biotechnol. 43, 1960–1965 (2025).
Google Scholar
Kim, A. et al. Identifying independent causal cell types for human diseases and risk variants. Cell Genom. https://doi.org/10.1016/j.xgen.2026.101325 (2026).
Landrum, M. J. et al. ClinVar: public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res. 42, D980–D985 (2014).
Google Scholar
Rentzsch, P., Witten, D., Cooper, G. M., Shendure, J. & Kircher, M. CADD: predicting the deleteriousness of variants throughout the human genome. Nucleic Acids Res. 47, D886–D894 (2019).
Google Scholar
Rives, A. et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc. Natl Acad. Sci. USA 118, e2016239118 (2021).
Google Scholar
Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023).
Google Scholar
Hayes, T. et al. Simulating 500 million years of evolution with a language model. Science 387, 850858 (2025).
Google Scholar
Tate, J. G. et al. COSMIC: the catalogue of somatic mutations in cancer. Nucleic Acids Res. 47, D941–D947 (2019).
Google Scholar
Chen, S. et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature 625, 92–100 (2024).
Google Scholar
Notin, P. et al. ProteinGym: large-scale benchmarks for protein fitness prediction and design. Adv. Neural Info. Process. Syst. 36, 64331–64379 (2023).
Google Scholar
Cheng, J. et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 381, eadg7492 (2023).
Google Scholar
Gao, H. et al. The landscape of tolerated genetic variation in humans and primates. Science 380, eabn8153 (2023).
Google Scholar
Ghosh, R. et al. Updated recommendation for the benign stand-alone ACMG/AMP criterion. Hum. Mutat. 39, 1525–1530 (2018).
Google Scholar
Avsec, Ž et al. Effective gene expression prediction from sequence by integrating long-range interactions. Nat. Methods 18, 1196–1203 (2021).
Google Scholar
Linder, J., Srivastava, D., Yuan, H., Agarwal, V. & Kelley, D. R. Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation. Nat. Genet. 57, 949–961 (2025).
Google Scholar
Avsec, Ž et al. Advancing regulatory variant effect prediction with AlphaGenome. Nature 649, 1206–1218 (2026).
Google Scholar
Amberger, J. S., Bocchini, C. A., Schiettecatte, F., Scott, A. F. & Hamosh, A. OMIM.org: Online Mendelian Inheritance in Man (OMIM), an online catalog of human genes and genetic disorders. Nucleic Acids Res. 43, D789–D798 (2015).
Google Scholar
Stenson, P. D. et al. The Human Gene Mutation Database (HGMD): optimizing its use in a clinical diagnostic or research setting. Hum. Genet. 139, 1197–1207 (2020).
Google Scholar
Jaganathan, K. et al. Predicting expression-altering promoter mutations with deep learning. Science 389, eads7373 (2025).
Google Scholar
Tomaz da Silva, P. et al. Nucleotide dependency analysis of genomic language models detects functional elements. Nat. Genet. 57, 2589–2602 (2025).
Google Scholar
Kanai, M. et al. Insights from complex trait fine-mapping across diverse populations. Preprint at medRxiv https://doi.org/10.1101/2021.09.03.21262975 (2021).
Bomba, L., Walter, K. & Soranzo, N. The impact of rare and low-frequency genetic variants in common disease. Genome Biol. 18, 77 (2017).
Google Scholar
Lee, S., Abecasis, G. R., Boehnke, M. & Lin, X. Rare-variant association analysis: study designs and statistical tests. Am. J. Hum. Genet. 95, 5–23 (2014).
Google Scholar
Backman, J. D. et al. Exome sequencing and analysis of 454,787 UK Biobank participants. Nature 599, 628–634 (2021).
Google Scholar
Karczewski, K. J. et al. Systematic single-variant and gene-based association testing of thousands of phenotypes in 394,841 UK Biobank exomes. Cell Genom. 2, 100168 (2022).
Google Scholar
Zhou, J. & Troyanskaya, O. G. Predicting effects of noncoding variants with deep learning–based sequence model. Nat. Methods 12, 931–934 (2015).
Google Scholar
Finucane, H. K. et al. Partitioning heritability by functional annotation using genome-wide association summary statistics. Nat. Genet. 47, 1228–1235 (2015).
Google Scholar
Weissbrod, O. et al. Functionally informed fine-mapping and polygenic localization of complex trait heritability. Nat. Genet. 52, 1355–1363 (2020).
Google Scholar
Márquez-Luna, C. et al. Incorporating functional priors improves polygenic prediction accuracy in UK Biobank and 23andMe data sets. Nat. Commun. 12, 6052 (2021).
Google Scholar
O’Connor, L. J. & Sella, G. Principled measures and estimates of trait polygenicity. Preprint at bioRxiv https://doi.org/10.1101/2025.07.10.664154 (2025).
Karollus, A., Mauermeier, T. & Gagneur, J. Current sequence-based models capture gene expression determinants in promoters but mostly ignore distal enhancers. Genome Biol. 24, 56 (2023).
Google Scholar
Fabiha, T. et al. A consensus variant-to-function score to functionally prioritize variants for disease. Preprint at bioRxiv https://doi.org/10.1101/2024.11.07.622307 (2024).
Finucane, H. K. et al. Heritability enrichment of specifically expressed genes identifies disease-relevant tissues and cell types. Nat. Genet. 50, 621–629 (2018).
Google Scholar
Zhang, Z. et al. Protein language models learn evolutionary statistics of interacting sequence motifs. Proc. Natl Acad. Sci. USA 121, e2406285121 (2024).
Google Scholar
Song, B., Buckler, E. S. & Stitzer, M. C. New whole-genome alignment tools are needed for tapping into plant diversity. Trends Plant Sci. 29, 355–369 (2024).
Google Scholar
Öztürk-Çolak, A. et al. FlyBase: updates to the Drosophila genes and genomes database. Genetics 227, iyad211 (2024).
Google Scholar
Qin, Z. et al. Genomic identification and functional characterization of essential genes in Caenorhabditis elegans. Genes Genomes Genet. 8, 981–997 (2018).
Google Scholar
Small, S., Blair, A. & Levine, M. Regulation of even-skipped stripe 2 in the Drosophila embryo. EMBO J. 11, 4047–4057 (1992).
Google Scholar
Wong, E. S. et al. Deep conservation of the enhancer regulatory code in animals. Science 370, eaax8137 (2020).
Google Scholar
Lewin, H. A. et al. Earth BioGenome project: sequencing life for the future of life. Proc. Natl Acad. Sci. USA 115, 4325–4333 (2018).
Google Scholar
Seplyarskiy, V. et al. A mutation rate model at the basepair resolution identifies the mutagenic effect of polymerase III transcription. Nat. Genet. 55, 2235–2242 (2023).
Google Scholar
Ye, C., Benegas, G., Albors, C., Li, J. C. & Song, Y. S. GPN-Star model source code. Zenodo https://doi.org/10.5281/zenodo.21501177 (2026).
Verbeek, M. M. et al. Mutations in the cyclic adenosine monophosphate response element of the tyrosine hydroxylase gene. Ann. Neurol. 62, 422–426 (2007).
Google Scholar
Kircher, M. et al. Saturation mutagenesis of twenty disease-associated regulatory elements at single base-pair resolution. Nat. Commun. 10, 3583 (2019).
Google Scholar
Bennett, M. K., Ngo, T. T., Athanikar, J. N., Rosenfeld, J. M. & Osborne, T. F. Co-stimulation of promoter for low density lipoprotein receptor gene by sterol regulatory element-binding protein and Sp1 is specifically disrupted by the yin yang 1 protein. J. Biol. Chem. 274, 13025–13032 (1999).
Google Scholar
Keep following us for the latest insights.
















