4 papers
HumanOmni-Speaker: Identifying Who said What and When
Detao Bai, Zhiheng Ma, Xihan Wei
While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, mult…
OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder
Detao Bai, Shimin Yao, Weixuan Chen +4
Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specifi…
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
Qize Yang, Shimin Yao, Weixuan Chen +7
With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands d…
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Jiaxing Zhao, Qize Yang, Yixing Peng +8
In human-centric scenes, the ability to simultaneously understand visual and auditory information is crucial. While recent omni models can process multiple modalities, they general…