2 papers
cs.LG2026
AuRA: Internalizing Audio Understanding into LLMs as LoRA
Bo Cheng, Lei Shi, Zhanyu Ma +5
Recent efforts to extend large language models (LLMs) to speech inputs typically rely on cascaded ASR-LLM pipelines, end-to-end speech-language models, or bridge/distillation-based…
cs.CV2026
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
Zhixuan Wu, Quanxing Zha, Teng Wang +6
Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively…