10 papers
HumanOmni-Speaker: Identifying Who said What and When
Detao Bai, Zhiheng Ma, Xihan Wei
While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, mult…
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
Boyuan Sun, Bowen Yin, Yuanming Li +2
We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts…
OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder
Detao Bai, Shimin Yao, Weixuan Chen +4
Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specifi…
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
Boyuan Sun, Jiaxing Zhao, Xiang Chen +2
In this paper, we introduce LLaVA-Octopus, a novel video multimodal large language model. LLaVA-Octopus adaptively weights features from different visual projectors based on user i…
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
Boyuan Sun, Jiaxing Zhao, Xihan Wei +1
In this paper, we present LLaVA-Scissor, a training-free token compression strategy designed for video multimodal large language models. Previous methods mostly attempt to compress…
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
Qize Yang, Shimin Yao, Weixuan Chen +7
With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands d…