collaborators

8 papers

cs.LG2026

PACT: Preserving Anchored Cores in Task-vectors for Model Merging

Ningyuan Shi, Zhipeng Zhou, Hao Wang +2

Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a single multi-task model. Most exi…

cs.LG2026

Delve into the Applicability of Advanced Optimizers for Multi-Task Learning

Zhipeng Zhou, Linxiao Cao, Pengcheng Wu +2

Multi-Task Learning (MTL) is a foundational machine learning problem that has seen extensive development over the past decade. Recently, various optimization-based MTL approaches h…

cs.AI2026

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory

Zhifei Xie, Zongzheng Hu, Fangda Ye +10

Proactivity is a core expectation for AGI. Prior work remains largely confined to laboratory settings, leaving a clear gap in real-world proactive agent: depth, complexity, ambigui…

cs.CL2025

Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models

Zhifei Xie, Ziyang Ma, Zihang Liu +7

Reasoning is essential for effective communication and decision-making. While recent advances in LLMs and MLLMs have shown that incorporating explicit reasoning significantly impro…

cs.SD2025

Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models

Zhifei Xie, Mingbao Lin, Zihang Liu +3

Recent advancements in multimodal reasoning have largely overlooked the audio modality. We introduce Audio-Reasoner, a large-scale audio language model for deep reasoning in audio…

cs.LG2025

Adapt in the Wild: Test-Time Entropy Minimization with Sharpness and Feature Regularization

Shuaicheng Niu, Guohao Chen, Deyu Chen +7

Test-time adaptation (TTA) may fail to improve or even harm the model performance when test data have: 1) mixed distribution shifts, 2) small batch sizes, 3) online imbalanced labe…