5 papers
scMRDR: A scalable and flexible framework for unpaired single-cell multi-omics data integration
Jianle Sun, Chaoqi Liang, Ran Wei +5
Advances in single-cell sequencing have enabled high-resolution profiling of diverse molecular modalities, while integrating unpaired multi-omics single-cell data remains challengi…
Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNA
Lifeng Qiao, Peng Ye, Yuchen Ren +5
Foundation models have made significant strides in understanding the genomic language of DNA sequences. However, previous models typically adopt the tokenization methods designed f…
IMWA: Iterative Model Weight Averaging Benefits Class-Imbalanced Learning Tasks
Zitong Huang, Ze Chen, Bowen Dong +3
Model Weight Averaging (MWA) is a technique that seeks to enhance model's performance by averaging the weights of multiple trained models. This paper first empirically finds that 1…
Toward Understanding BERT-Like Pre-Training for DNA Foundation Models
Chaoqi Liang, Lifeng Qiao, Peng Ye +9
With the success of large-scale pre-training in language tasks, there is an increasing trend of applying it to the domain of life sciences. In particular, pre-training methods base…
LayerMatch: Do Pseudo-labels Benefit All Layers?
Chaoqi Liang, Guanglei Yang, Lifeng Qiao +4
Deep neural networks have achieved remarkable performance across various tasks when supplied with large-scale labeled data. However, the collection of labeled data can be time-cons…