collaborators

11 papers

cs.LG2026

Closed-Form Spectral Regularization for Multi-Task Model Merging

Yongxian Wei, Runxi Cheng, Xingxuan Zhang +4

Model merging combines several independently fine-tuned experts into a single multi-task model without any training data, reducing the storage, serving, and decentralized-developme…

cs.CL2026

Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory

Runxi Cheng, Yuchen Guan, Yongxian Wei +7

Scaling conditional memory offers a promising way to increase language-model capacity, but existing methods such as Engram learn large memory tables from scratch during pre-trainin…

cs.CL2026

Spectral Characterization and Mitigation of Sequential Knowledge Editing Collapse

Chi Zhang, Mengqi Zhang, Xiaotian Ye +5

Sequential knowledge editing in large language models often causes catastrophic collapse of the model's general abilities, especially for parameter-modifying methods. Existing appr…

cs.AI2026

Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis

Yongxian Wei, Yilin Zhao, Zixuan Hu +7

Data synthesis for training large reasoning models offers a scalable alternative to limited, human-curated datasets, enabling the creation of high-quality data. However, existing a…

cs.AI2026

OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging

Yongxian Wei, Runxi Cheng, Weike Jin +7

Foundation models update slowly due to resource-intensive training, whereas domain-specific models evolve rapidly between releases. Model merging seeks to combine multiple expert m…

cs.CV2025

GSRender: Deduplicated Occupancy Prediction via Weakly Supervised 3D Gaussian Splatting

Qianpu Sun, Changyong Shu, Sifan Zhou +6

Weakly-supervised 3D occupancy perception is crucial for vision-based autonomous driving in outdoor environments. Previous methods based on NeRF often face a challenge in balancing…