1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2026
The Bridge-Garden Dilemma in LLM Distillation: Why Mixing Hard and Soft Labels Works
Guanghui Wang, Kaiwen Lv Kacuila, Zhiyong Yang +5
Knowledge distillation (KD) transfers knowledge from a large teacher model to a smaller student. In language modeling, the student is trained either on tokens sampled from the teac…
stat.ML2026
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
Kuo-Wei Lai, Guanghui Wang, Molei Tao +1
Overparameterized ML models, including neural networks, typically induce underdetermined training objectives with multiple global minima. The implicit bias refers to the limiting g…
cs.CV2024★ 1 cited
SuperLoRA: Parameter-Efficient Unified Adaptation of Multi-Layer Attention Modules
Xiangyu Chen, Jing Liu, Ye Wang +4
Low-rank adaptation (LoRA) and its variants are widely employed in fine-tuning large models, including large language models for natural language processing and diffusion models fo…