collaborators

10 papers

cs.CV2026

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models

Zihan Song, Shuo Ye, Bo Zhao +4

Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindranc…

cs.CV2026

Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning

Yixin Ji, Fanghua Ye, Juntao Li +5

Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing…

eess.AS2026

StepAudio 2.5 Technical Report

Bin Lin, Bo Zhao, Boyong Wu +98

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…

cs.CV2026

AffectVerse: Emotional World Models for Multimodal Affective Computing

Bo Zhao, Fanghua Ye, Yixin Ji +3

Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large language models (MLLMs), however, o…

cs.CV2026

UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings

Bo Zhao, Maosheng Pang, Chen Zhang +3

Existing works typically focus on presentation generation under isolated input settings, whereas real-world use cases span diverse scenarios, including vague user prompts, long doc…

cs.LG2026

Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition

Zeheng Wang, Bo Zhao, Yijie Zhu +6

Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large language models excel at multimodal…