A Phylogenetic Approach to Genomic Language Modeling
By: Carlos Albors , Jianan Canal Li , Gonzalo Benegas and more
Potential Business Impact:
Finds important parts of our DNA.
Genomic language models (gLMs) have shown mostly modest success in identifying evolutionarily constrained elements in mammalian genomes. To address this issue, we introduce a novel framework for training gLMs that explicitly models nucleotide evolution on phylogenetic trees using multispecies whole-genome alignments. Our approach integrates an alignment into the loss function during training but does not require it for making predictions, thereby enhancing the model's applicability. We applied this framework to train PhyloGPN, a model that excels at predicting functionally disruptive variants from a single sequence alone and demonstrates strong transfer learning capabilities.
Similar Papers
Evaluating DNA function understanding in genomic language models using evolutionarily implausible sequences
Quantitative Methods
Helps computers design new DNA that works.
Open-weight genome language model safeguards: Assessing robustness via adversarial fine-tuning
Machine Learning (CS)
Computers can still make dangerous viruses, even if hidden.
CodonMoE: DNA Language Models for mRNA Analyses
Genomics
Makes DNA computers understand RNA code better.