Publications (12)
HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation
Feng Xiong, Hongling Xu, Yifei Wang +3
Self-taught reasoners (STaRs) enhance the mathematical reasoning abilities of large language models (LLMs) by leveraging self-generated responses for self-training. Recent studies…
Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis
Yongxian Wei, Yilin Zhao, Zixuan Hu +7
Data synthesis for training large reasoning models offers a scalable alternative to limited, human-curated datasets, enabling the creation of high-quality data. However, existing a…
Learn To Learn More Precisely
Runxi Cheng, Yongxian Wei, Xianglong He +5
Meta-learning has been extensively applied in the domains of few-shot learning and fast adaptation, achieving remarkable performance. While Meta-learning methods like Model-Agnosti…
Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory
Runxi Cheng, Yuchen Guan, Yongxian Wei +7
Scaling conditional memory offers a promising way to increase language-model capacity, but existing methods such as Engram learn large memory tables from scratch during pre-trainin…
Enhancing Logits Distillation with Plug\&Play Kendall's Ranking Loss
Yuchen Guan, Runxi Cheng, Kang Liu +1
Knowledge distillation typically minimizes the Kullback-Leibler (KL) divergence between teacher and student logits. However, optimizing the KL divergence can be challenging for the…
Spectral Characterization and Mitigation of Sequential Knowledge Editing Collapse
Chi Zhang, Mengqi Zhang, Xiaotian Ye +5
Sequential knowledge editing in large language models often causes catastrophic collapse of the model's general abilities, especially for parameter-modifying methods. Existing appr…
Mixture of Neuron Experts
Runxi Cheng, Yuchen Guan, Yucheng Ding +6
In this work, we first explore whether the parameters activated by the MoE layer remain highly sparse at inference. We perform a sparsification study on several representative MoE…
Multi-Task Model Merging via Adaptive Weight Disentanglement
Feng Xiong, Runxi Cheng, Wang Chen +4
Model merging has recently gained attention as an economical and scalable approach to incorporate task-specific weights from various tasks into a unified multi-task model. For exam…
GSRender: Deduplicated Occupancy Prediction via Weakly Supervised 3D Gaussian Splatting
Qianpu Sun, Changyong Shu, Sifan Zhou +6
Weakly-supervised 3D occupancy perception is crucial for vision-based autonomous driving in outdoor environments. Previous methods based on NeRF often face a challenge in balancing…
Closed-Form Spectral Regularization for Multi-Task Model Merging
Yongxian Wei, Runxi Cheng, Xingxuan Zhang +4
Model merging combines several independently fine-tuned experts into a single multi-task model without any training data, reducing the storage, serving, and decentralized-developme…
Whoever Started the Interference Should End It: Guiding Data-Free Model Merging via Task Vectors
Runxi Cheng, Feng Xiong, Yongxian Wei +2
Model merging seeks to integrate task-specific expert models into a unified architecture while preserving multi-task generalization capabilities, yet parameter interference between…
OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging
Yongxian Wei, Runxi Cheng, Weike Jin +7
Foundation models update slowly due to resource-intensive training, whereas domain-specific models evolve rapidly between releases. Model merging seeks to combine multiple expert m…