3 papers
cs.CV2026
A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training
Kaichen Li, Zhilin Zhu, Jianhao Huang +7
In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Exi…
cs.SD2026
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
Haowei Lou, Hye-Young Paik, Dai Jia +2
Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for spe…
cs.AI2026
AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models
Zibo Shao, Baochen Xiong, Chengdong Xu +6
Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are…