Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Wanda++: Pruning Large Language Models via Regional Gradients
Yifan Yang, Kai Zhen, Bhavana Ganesh +11
Large Language Models (LLMs) pruning seeks to remove unimportant weights for inference speedup with minimal accuracy impact. However, existing methods often suffer from accuracy de…
cs.LG2024
SeRA: Self-Reviewing and Alignment of Large Language Models using Implicit Reward Margins
Jongwoo Ko, Saket Dingliwal, Bhavana Ganesh +3
Direct alignment algorithms (DAAs), such as direct preference optimization (DPO), have become popular alternatives for Reinforcement Learning from Human Feedback (RLHF) due to thei…