From the 1 of 12 linked papers with an AI index.
12 papers
MMAG: A Multi-Control Mixed Audio Generation Benchmark
Zihao Zheng, Xuenan Xu, Jiahao Mei +5
Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluat…
SLVMBench: Skill Learning from Video Memory
Yudong Yang, Guangzhi Sun, Yixuan Li +1
The paper presents SLVMBench, a benchmark that tests whether video large language models can learn procedural skills from multi‑hour video streams and then use that knowledge to an…
video-SALMONN-R: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding
Yixuan Li, Guangzhi Sun, Yudong Yang +1
Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to…
OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs
Guangzhi Sun, Yixuan Li, Yudong Yang +1
Audio-visual large language models (LLMs) hold strong promise for long-form video understanding, yet their long-video inference is fundamentally limited by the linear growth of vid…
video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
Guangzhi Sun, Yixuan Li, Xiaodong Wu +4
Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhance…
OCR-Enhanced Multimodal ASR Can Read While Listening
Junli Chen, Changli Tang, Yixuan Li +2
Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to…