activity
20192026
most citedSALMONN: Towards Generic Hearing Abilities for Large Language Models

20 citations · 82 across the 51 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

SLVMBench: Skill Learning from Video Memory

Yudong Yang, Guangzhi Sun, Yixuan Li +1

We introduce Skill Learning from Video Memory (SLVMBench), the first benchmark that jointly evaluates whether video large language models (video-LLMs) can learn skills from long vi…

cs.CV2025

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

Guangzhi Sun, Yixuan Li, Xiaodong Wu +4

Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhance…

cs.CV2025★ 1 cited

video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models

Changli Tang, Yixuan Li, Yudong Yang +5

We present video-SALMONN 2, a family of audio-visual large language models that set new state-of-the-art (SOTA) results in video description and question answering (QA). Our core c…

cs.CV2025

Audio-centric Video Understanding Benchmark without Text Shortcut

Yudong Yang, Jimin Zhuang, Guangzhi Sun +7

Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information.…

cs.CV2025

Improving LLM Video Understanding with 16 Frames Per Second

Yixuan Li, Changli Tang, Jimin Zhuang +5

Human vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted f…

cs.CV2025★ 2 cited

video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model

Guangzhi Sun, Yudong Yang, Jimin Zhuang +5

While recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been li…