5 papers · 1 filter
BLADE: Scalable Bi-level Adaptive Data Selection for LLM Training
Jiaxing Wang, Deping Xiang, Jin Xu +9
As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive lear…
TANDEM: Bi-Level Data Mixture Optimization with Twin Networks
Jiaxing Wang, Deping Xiang, Jin Xu +9
The capabilities of large language models (LLMs) significantly depend on training data drawn from various domains. Optimizing domain-specific mixture ratios can be modeled as a bi-…
Spectral Disentanglement and Enhancement: A Dual-domain Contrastive Framework for Representation Learning
Jinjin Guo, Yexin Li, Zhichao Huang +5
Large-scale multimodal contrastive learning has recently achieved impressive success in learning rich and transferable representations, yet it remains fundamentally limited by the…
The Primacy of Magnitude in Low-Rank Adaptation
Zicheng Zhang, Haoran Li, Yifeng Zhang +5
Low-Rank Adaptation (LoRA) offers a parameter-efficient paradigm for tuning large models. While recent spectral initialization methods improve convergence and performance over the…
FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
Jiaoyang Li, Jun Fang, Tianhao Gao +5
Representation learning is fundamental to modern machine learning, powering applications such as text retrieval and multimodal understanding. However, learning robust and generaliz…