2 papers
cs.CL2026
Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction
Qi Chen, Yunfei Chu, Haolin He +15
Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed tex…
cs.SD2026
LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension
Wen Huang, Yunfei Chu, Meng Gao +2
General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal…