3 papers
cs.AI2026
Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
Zhi Zeng, Cheng Zhang, Zesheng Yang +9
Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models…
cs.CV2025
Seed1.5-VL Technical Report
Dong Guo, Faming Wu, Feida Zhu +194
We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…
cs.LG2025
From Layers to States: A State Space Model Perspective to Deep Neural Network Layer Dynamics
Qinshuo Liu, Weiqin Zhao, Wei Huang +3
The depth of neural networks is a critical factor for their capability, with deeper models often demonstrating superior performance. Motivated by this, significant efforts have bee…