collaborators

5 papers

cs.CV2026

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models

Zihan Song, Shuo Ye, Bo Zhao +4

Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindranc…

cs.LG2026

Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition

Zeheng Wang, Bo Zhao, Yijie Zhu +6

Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large language models excel at multimodal…

cs.CV2026

Complementarity-Supervised Spectral-Band Routing for Multimodal Emotion Recognition

Zhexian Huang, Bo Zhao, Hui Ma +5

Multimodal emotion recognition fuses cues such as text, video, and audio to understand individual emotional states. Prior methods face two main limitations: mechanically relying on…

cs.CV2025

Diff-Palm: Realistic Palmprint Generation with Polynomial Creases and Intra-Class Variation Controllable Diffusion Models

Jianlong Jin, Chenglong Zhao, Ruixin Zhang +8

Palmprint recognition is significantly limited by the lack of large-scale publicly available datasets. Previous methods have adopted Bézier curves to simulate the palm creases, wh…

cs.CV2025

PVTree: Realistic and Controllable Palm Vein Generation for Recognition Tasks

Sheng Shang, Chenglong Zhao, Ruixin Zhang +7

Palm vein recognition is an emerging biometric technology that offers enhanced security and privacy. However, acquiring sufficient palm vein data for training deep learning-based r…