2 citations · 4 across the 17 of their papers we have counts for
18 papers · 1 filter
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
Boyuan Sun, Bowen Yin, Yuanming Li +2
We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts…
OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder
Detao Bai, Shimin Yao, Weixuan Chen +4
Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specifi…
HumanOmni-Speaker: Identifying Who said What and When
Detao Bai, Zhiheng Ma, Xihan Wei
While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, mult…
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
Yuan-Ming Li, Qize Yang, Nan Lei +5
Recent advances in motion-aware large language models have shown remarkable promise for jointly learning motion understanding and generation knowledge. However, these models typica…
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
Shenghao Fu, Qize Yang, Yuan-Ming Li +3
Long video understanding is still challenging for recent Large Video-Language Models (LVLMs) due to the conflict between long-form temporal understanding and detailed spatial perce…
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
Boyuan Sun, Jiaxing Zhao, Xihan Wei +1
In this paper, we present LLaVA-Scissor, a training-free token compression strategy designed for video multimodal large language models. Previous methods mostly attempt to compress…