2 papers
cs.MM2026
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models
Yuwen Wang, Tian-Hao Zhang, Minghao Cai +7
Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rath…
cs.CV2026
One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models
Sudharshan Balaji, Yili Ren, Guangjing Wang +2
Machine unlearning is widely used to remove hazardous knowledge from large language models. Modern Vision-Language Models (VLMs), however, process both text and visual inputs, rais…