3 papers
cs.CL2026
STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Models
Xinhao Sun, Huaijin Zhao, Maoliang Li +4
Diffusion Language Models (DLMs) enable parallel decoding via iterative denoising, where remasking strategies play a critical role in balancing inference speed and output quality.…
cs.LG2025
MIN-Merging: Merge the Important Neurons for Model Merging
Yunfei Liang
Recent advances in deep learning have led to a surge of open-source models across diverse domains. While model merging offers a promising way to combine their strengths, existing a…
cs.LG2025
FedHQ: Hybrid Runtime Quantization for Federated Learning
Zihao Zheng, Ziyao Wang, Xiuping Cui +6
Federated Learning (FL) is a decentralized model training approach that preserves data privacy but struggles with low efficiency. Quantization, a powerful training optimization tec…