4 papers
PrunePath: Towards Highly Structured Sparse Language Models
Zhexuan Gu, Zixun Fu, Yancheng Yuan
Feed-forward networks (FFNs) dominate the parameter count and computation of modern language models, yet existing pruning methods often struggle to convert sparsity into hardware-f…
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
Zonglin Yang, Zhexuan Gu, Yancheng Yuan
Training a deep neural network with the outputs of selected layers satisfying linear constraints is required in many contemporary data-driven applications. While this can be achiev…
Accelerating RLHF Training with Reward Variance Increase
Zonglin Yang, Zhexuan Gu, Houduo Qi +1
Reinforcement learning from human feedback (RLHF) is an essential technique for ensuring that large language models (LLMs) are aligned with human values and preferences during the…
HOT: An Efficient Halpern Accelerating Algorithm for Optimal Transport Problems
Guojun Zhang, Zhexuan Gu, Yancheng Yuan +1
This paper proposes an efficient HOT algorithm for solving the optimal transport (OT) problems with finite supports. We particularly focus on an efficient implementation of the HOT…