From the 1 of 17 linked papers with an AI index.
17 papers
SLVMBench: Skill Learning from Video Memory
Yudong Yang, Guangzhi Sun, Yixuan Li +1
The paper presents SLVMBench, a benchmark that tests whether video large language models can learn procedural skills from multi‑hour video streams and then use that knowledge to an…
video-SALMONN-R: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding
Yixuan Li, Guangzhi Sun, Yudong Yang +1
Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to…
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments
Zhan Liu, Changli Tang, Yuxin Wang +7
Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundame…
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
Changli Tang, Qinfan Xiao, Ke Mei +3
While embeddings from multimodal large language models (LLMs) excel as general-purpose representations, their application to dynamic modalities like audio and video remains underex…
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
Changli Tang, Tianyi Wang, Fengyun Rao +2
Spoken dialogue is a primary source of information in videos; therefore, accurately identifying who spoke what and when is essential for deep video understanding. We introduce D-OR…
video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
Guangzhi Sun, Yixuan Li, Xiaodong Wu +4
Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhance…