8 papers
AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
Yuan Zhang, Chun-Kai Fan, Sicheng Yu +6
Inspired by text prompts in large language models, visual prompts have been explored to enhance the perceptual capabilities of large vision-language models (LVLMs). However, perfor…
Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control
Zhe Li, Cheng Chi, Yangyang Wei +9
Humans intuitively move to sound, but current humanoid robots lack expressive improvisational capabilities, confined to predefined motions or sparse commands. Generating motion fro…
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
Zhe Li, Cheng Chi, Boan Zhu +11
Humans learn locomotion through visual observation, interpreting visual content first before imitating actions. However, state-of-the-art humanoid locomotion systems rely on either…
ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
Yuan Zhang, Ming Lu, Junwen Pan +4
Recent advances in multimodal reasoning models have demonstrated impressive capabilities across text and vision. However, even leading models exhibit redundant self-reflection when…
From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
Zhe Li, Cheng Chi, Yangyang Wei +7
Natural language offers a natural interface for humanoid robots, but existing language-guided humanoid locomotion pipelines remain cumbersome and untrustworthy. They typically deco…
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Yuan Zhang, Chun-Kai Fan, Junpeng Ma +8
In vision-language models (VLMs), visual tokens usually bear a significant amount of computational overhead despite sparsity of information in them when compared to text tokens. To…