10 papers
PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
Zihan Song, Shuo Ye, Bo Zhao +4
Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindranc…
Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning
Yixin Ji, Fanghua Ye, Juntao Li +5
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing…
StepAudio 2.5 Technical Report
Bin Lin, Bo Zhao, Boyong Wu +98
Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…
AffectVerse: Emotional World Models for Multimodal Affective Computing
Bo Zhao, Fanghua Ye, Yixin Ji +3
Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large language models (MLLMs), however, o…
UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings
Bo Zhao, Maosheng Pang, Chen Zhang +3
Existing works typically focus on presentation generation under isolated input settings, whereas real-world use cases span diverse scenarios, including vague user prompts, long doc…
Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition
Zeheng Wang, Bo Zhao, Yijie Zhu +6
Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large language models excel at multimodal…