collaborators

6 papers

cs.CV2026

HumanOmni-Speaker: Identifying Who said What and When

Detao Bai, Zhiheng Ma, Xihan Wei

While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, mult…

cs.CV2026

OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

Detao Bai, Shimin Yao, Weixuan Chen +4

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specifi…

cs.CV2025

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Qize Yang, Shimin Yao, Weixuan Chen +7

With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands d…

cs.SD2025

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization

Detao Bai, Zhiheng Ma, Xihan Wei +1

The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tas…

cs.CV2025

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Jiaxing Zhao, Qize Yang, Yixing Peng +8

In human-centric scenes, the ability to simultaneously understand visual and auditory information is crucial. While recent omni models can process multiple modalities, they general…

cs.CV2025

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis

Qize Yang, Detao Bai, Yi-Xing Peng +1

Understanding emotions accurately is essential for fields like human-computer interaction. Due to the complexity of emotions and their multi-modal nature (e.g., emotions are influe…