3 papers
cs.CV2026
Audio-Visual Intelligence in Large Foundation Models
You Qin, Kai Liu, Shengqiong Wu +12
Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate…
cs.AI2026
Synergizing Large Language Models and Task-specific Models for Time Series Anomaly Detection
Feiyi Chen, Leilei Zhang, Guansong Pang +2
In anomaly detection, methods based on large language models (LLMs) can incorporate expert knowledge by reading professional document, while task-specific small models excel at ext…
cs.CV2025
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
Yue Zhang, Yingzhao Jian, Hehe Fan +2
Recent advancements in multimodal large language models (MLLMs) have demonstrated considerable potential for comprehensive 3D scene understanding. However, existing approaches typi…