23 papers
LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning
Hao Jiang, Enneng Yang, Guojie Zhu +7
Continual learning capability is critical for Industrial LLMs, as deployed models must be continuously updated to meet evolving requirements and environments, rather than repeatedl…
Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation
Chu Zhao, Enneng Yang, Jianzhe Zhao +1
Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior distributions by minimizing preference al…
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
Chu Zhao, Enneng Yang, Yuting Liu +2
Test-time reinforcement learning generates multiple candidate answers via repeated rollouts and performs online updates using pseudo-labels constructed by majority voting. To reduc…
Automatic Pruning Discovery for Large Language Models
Haidong Kang, Lihong Lin, Enneng Yang +2
Large language models (LLMs) have achieved remarkable performance on a wide range of tasks, hindering real-world deployment due to their massive size. Existing pruning methods (e.g…
OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging
Yongxian Wei, Runxi Cheng, Weike Jin +7
Foundation models update slowly due to resource-intensive training, whereas domain-specific models evolve rapidly between releases. Model merging seeks to combine multiple expert m…
Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
Yingsha Xie, Tiansheng Huang, Enneng Yang +5
Safety alignment incurs safety tax that perturbs a large reasoning model's (LRM) general reasoning ability. Existing datasets used for safety alignment for an LRM are usually const…