8 papers
SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
Xiangchao Yan, Runjian Chen, Bo Zhang +11
Annotating 3D LiDAR point clouds for perception tasks is fundamental for many applications e.g., autonomous driving, yet it still remains notoriously labor-intensive. Pretraining-f…
SE-Merging: A Self-Enhanced Approach for Dynamic Model Merging
Zijun Chen, Zhanpeng Zhou, Bo Zhang +3
Model merging has gained increasing attention due to its intriguing property: interpolating the parameters of different task-specific fine-tuned models leads to multi-task abilitie…
ProtoReasoning: Prototypes as the Foundation for Generalizable Reasoning in LLMs
Feng He, Zijun Chen, Xinnian Liang +4
Recent advances in Large Reasoning Models (LRMs) trained with Long Chain-of-Thought (Long CoT) reasoning have demonstrated remarkable cross-domain generalization capabilities. Howe…
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
Jinbo Wang, Mingze Wang, Zhanpeng Zhou +3
Transformers consist of diverse building blocks, such as embedding layers, normalization layers, self-attention mechanisms, and point-wise feedforward networks. Thus, understanding…
On the Role of Label Noise in the Feature Learning Process
Andi Han, Wei Huang, Zhanpeng Zhou +5
Deep learning with noisy labels presents significant challenges. In this work, we theoretically characterize the role of label noise from a feature learning perspective. Specifical…
New Evidence of the Two-Phase Learning Dynamics of Neural Networks
Zhanpeng Zhou, Yongyi Yang, Mahito Sugiyama +1
Understanding how deep neural networks learn remains a fundamental challenge in modern machine learning. A growing body of evidence suggests that training dynamics undergo a distin…