R.I.P.
π»
Ghosted
DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks
June 29, 2026 Β· Grace Period Β· + Add venue
Authors
Romain Karpinsky, Julien Mozziconacci, MickaΓ«l Delcey
arXiv ID
2606.30140
Category
q-bio.GN
Cross-listed
cs.CL
Citations
0
Abstract
Recent breakthroughs in foundation models and Large Language Models (LLMs) have introduced new opportunities for studying and decoding genomic sequences. Several state-of-the-art approaches, such as DNABERT2, rely on transformer-based architectures, while others, such as ConvNova, still build upon more conventional convolutional models. However, systematic benchmark comparisons across these methods remain scarce. Given that transformer-based models require extensive and costly pretraining, it is crucial to evaluate whether their performance gains justify this overhead. Moreover, LLMs such as DNABERT2 typically rely on Byte Pair Encoding (BPE) tokenization, whose relevance for DNA sequence representation is still debated within the genomics community. In this work, we investigate three key questions: (i) do transformer-based models provide sufficient improvements on fine-tuning tasks upon heavy pretraining, (ii) what is the actual contribution of pretraining in this setting, and (iii) how does BPE tokenization impact performance on genomics-related tasks?
Community Contributions
Found the code? Know the venue? Think something is wrong? Let us know!
π Similar Papers
In the same crypt β q-bio.GN
R.I.P.
π»
Ghosted
Accurate Genomic Prediction Of Human Height
R.I.P.
π»
Ghosted
Synergistic Drug Combination Prediction by Integrating Multi-omics Data in Deep Learning Models
π
π
Old Age
GateKeeper: A New Hardware Architecture for Accelerating Pre-Alignment in DNA Short Read Mapping
R.I.P.
π»
Ghosted
Tasks, Techniques, and Tools for Genomic Data Visualization
π
π
Old Age