works on

From the 1 of 12 linked papers with an AI index.

collaborators

12 papers

cs.SD2026

MMAG: A Multi-Control Mixed Audio Generation Benchmark

Zihao Zheng, Xuenan Xu, Jiahao Mei +5

Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluat…

cs.CV2026

SLVMBench: Skill Learning from Video Memory

Yudong Yang, Guangzhi Sun, Yixuan Li +1

The paper presents SLVMBench, a benchmark that tests whether video large language models can learn procedural skills from multi‑hour video streams and then use that knowledge to an…

cs.CV2026

video-SALMONN-R: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

Yixuan Li, Guangzhi Sun, Yudong Yang +1

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to…

cs.AI2026

OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs

Guangzhi Sun, Yixuan Li, Yudong Yang +1

Audio-visual large language models (LLMs) hold strong promise for long-form video understanding, yet their long-video inference is fundamentally limited by the linear growth of vid…

cs.CV2026

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

Guangzhi Sun, Yixuan Li, Xiaodong Wu +4

Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhance…

cs.SD2026

OCR-Enhanced Multimodal ASR Can Read While Listening

Junli Chen, Changli Tang, Yixuan Li +2

Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to…