1 citations · 1 across the 11 of their papers we have counts for
12 papers
Hybrid Policy Distillation for LLMs
Wenhong Zhu, Ruobing Xie, Rui Wang +1
Knowledge distillation (KD) is a powerful paradigm for compressing large language models (LLMs), whose effectiveness depends on intertwined choices of divergence direction, optimiz…
EtCon: Edit-then-Consolidate for Reliable Knowledge Editing
Ruilin Li, Yibin Wang, Wenhong Zhu +5
Knowledge editing aims to update specific facts in large language models (LLMs) without full retraining. Prior efforts sought to tune the knowledge layers of LLMs, achieving improv…
MrRoPE: Mixed-radix Rotary Position Embedding
Qingyuan Tian, Wenhong Zhu, Xiaoran Liu +2
Rotary Position Embedding (RoPE)-extension refers to modifying or generalizing the Rotary Position Embedding scheme to handle longer sequences than those encountered during pre-tra…
Proximal Supervised Fine-Tuning
Wenhong Zhu, Ruobing Xie, Rui Wang +3
Supervised fine-tuning (SFT) of foundation models often leads to poor generalization, where prior capabilities deteriorate after tuning on new tasks or domains. Inspired by trust-r…
Flexible Realignment of Language Models
Wenhong Zhu, Ruobing Xie, Weinan Zhang +1
Realignment becomes necessary when a language model (LM) fails to meet expected performance. We propose a flexible realignment framework that supports quantitative control of align…
Adding Alignment Control to Language Models
Wenhong Zhu, Weinan Zhang, Rui Wang
Post-training alignment has increasingly become a crucial factor in enhancing the usability of language models (LMs). However, the strength of alignment varies depending on individ…