collaborators

8 papers

cs.CV2025

SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations

Xiangchao Yan, Runjian Chen, Bo Zhang +11

Annotating 3D LiDAR point clouds for perception tasks is fundamental for many applications e.g., autonomous driving, yet it still remains notoriously labor-intensive. Pretraining-f…

cs.AI2025

SE-Merging: A Self-Enhanced Approach for Dynamic Model Merging

Zijun Chen, Zhanpeng Zhou, Bo Zhang +3

Model merging has gained increasing attention due to its intriguing property: interpolating the parameters of different task-specific fine-tuned models leads to multi-task abilitie…

cs.CL2025

ProtoReasoning: Prototypes as the Foundation for Generalizable Reasoning in LLMs

Feng He, Zijun Chen, Xinnian Liang +4

Recent advances in Large Reasoning Models (LRMs) trained with Long Chain-of-Thought (Long CoT) reasoning have demonstrated remarkable cross-domain generalization capabilities. Howe…

cs.LG2025

The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

Jinbo Wang, Mingze Wang, Zhanpeng Zhou +3

Transformers consist of diverse building blocks, such as embedding layers, normalization layers, self-attention mechanisms, and point-wise feedforward networks. Thus, understanding…

stat.ML2025

On the Role of Label Noise in the Feature Learning Process

Andi Han, Wei Huang, Zhanpeng Zhou +5

Deep learning with noisy labels presents significant challenges. In this work, we theoretically characterize the role of label noise from a feature learning perspective. Specifical…

cs.LG2025

New Evidence of the Two-Phase Learning Dynamics of Neural Networks

Zhanpeng Zhou, Yongyi Yang, Mahito Sugiyama +1

Understanding how deep neural networks learn remains a fundamental challenge in modern machine learning. A growing body of evidence suggests that training dynamics undergo a distin…