2 citations · 3 across the 9 of their papers we have counts for
8 papers · 1 filter
SLVMBench: Skill Learning from Video Memory
Yudong Yang, Guangzhi Sun, Yixuan Li +1
We introduce Skill Learning from Video Memory (SLVMBench), the first benchmark that jointly evaluates whether video large language models (video-LLMs) can learn skills from long vi…
video-SALMONN-R: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding
Yixuan Li, Guangzhi Sun, Yudong Yang +1
Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to…
video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
Guangzhi Sun, Yixuan Li, Xiaodong Wu +4
Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhance…
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
Changli Tang, Yixuan Li, Yudong Yang +5
We present video-SALMONN 2, a family of audio-visual large language models that set new state-of-the-art (SOTA) results in video description and question answering (QA). Our core c…
Audio-centric Video Understanding Benchmark without Text Shortcut
Yudong Yang, Jimin Zhuang, Guangzhi Sun +7
Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information.…
Improving LLM Video Understanding with 16 Frames Per Second
Yixuan Li, Changli Tang, Jimin Zhuang +5
Human vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted f…