7 papers
Reversible Diffusion Decoding for Diffusion Language Models
Xinyun Wang, Min Zhang, Sen Cui +4
Diffusion language models enable parallel token generation through block-wise decoding, but their irreversible commitments can lead to stagnation, where the reverse diffusion proce…
From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models
Zhikang Chen, Tingting Zhu
A world model is an AI system that simulates how an environment evolves under actions, enabling planning through imagined futures rather than reactive perception. Current world mod…
From Coefficients to Directions: Rethinking Model Merging with Directional Alignment
Zhikang Chen, Sen Cui, Deheng Ye +5
Model merging has emerged as a practical paradigm for integrating multiple independently trained models into a single model without joint retraining. Previous studies have demonstr…
Merging without Forgetting: Continual Fusion of Task-Specific Models via Optimal Transport
Zecheng Pan, Zhikang Chen, Ding Li +9
Merging models fine-tuned for different tasks into a single unified model has become an increasingly important direction for building versatile, efficient multi-task systems. Exist…
Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thought
Zhikang Chen, Sen Cui, Deheng Ye +3
Large Language Models (LLMs) have demonstrated strong reasoning capabilities through \emph{Chain-of-Thought} (CoT) prompting, which enables step-by-step intermediate reasoning. How…
Decentralized Dynamic Cooperation of Personalized Models for Federated Continual Learning
Danni Yang, Zhikang Chen, Sen Cui +6
Federated continual learning (FCL) has garnered increasing attention for its ability to support distributed computation in environments with evolving data distributions. However, t…