works on

From the 1 of 17 linked papers with an AI index.

activity
20242026
collaborators

17 papers

cs.CV2026

SLVMBench: Skill Learning from Video Memory

Yudong Yang, Guangzhi Sun, Yixuan Li +1

The paper presents SLVMBench, a benchmark that tests whether video large language models can learn procedural skills from multi‑hour video streams and then use that knowledge to an…

cs.CV2026

video-SALMONN-R: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

Yixuan Li, Guangzhi Sun, Yudong Yang +1

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to…

cs.CV2026

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

Zhan Liu, Changli Tang, Yuxin Wang +7

Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundame…

cs.CV2026

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

Changli Tang, Qinfan Xiao, Ke Mei +3

While embeddings from multimodal large language models (LLMs) excel as general-purpose representations, their application to dynamic modalities like audio and video remains underex…

cs.CV2026

D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning

Changli Tang, Tianyi Wang, Fengyun Rao +2

Spoken dialogue is a primary source of information in videos; therefore, accurately identifying who spoke what and when is essential for deep video understanding. We introduce D-OR…

cs.CV2026

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

Guangzhi Sun, Yixuan Li, Xiaodong Wu +4

Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhance…