4 papers
REAR: Test-time Preference Realignment through Reward Decomposition
Fuxiang Zhang, Pengcheng Wang, Chenran Li +6
Aligning large language models (LLMs) with diverse user preferences is a critical yet challenging task. While post-training methods can adapt models to specific needs, they often r…
AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations
Penghui Yang, Zhonghan Zhang, Yue Li +6
Density functional theory (DFT) serves as the basis for computational discovery in materials science and chemistry, yet each calculation demands extensive human effort: adjusting a…
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
Penghui Yang, Cunxiao Du, Fengzhuo Zhang +4
As Large Language Models (LLMs) can now process extremely long contexts, efficient inference over these extended inputs has become increasingly important, especially for emerging a…
Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary Head
Penghui Yang, Chen-Chen Zong, Sheng-Jun Huang +2
Traditional knowledge distillation focuses on aligning the student's predicted probabilities with both ground-truth labels and the teacher's predicted probabilities. However, the t…