activity
20242026
collaborators

8 papers

cs.CV2026

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs

Yuan Zhang, Chun-Kai Fan, Sicheng Yu +6

Inspired by text prompts in large language models, visual prompts have been explored to enhance the perceptual capabilities of large vision-language models (LVLMs). However, perfor…

cs.RO2026

Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control

Zhe Li, Cheng Chi, Yangyang Wei +9

Humans intuitively move to sound, but current humanoid robots lack expressive improvisational capabilities, confined to predefined motions or sparse commands. Generating motion fro…

cs.RO2026

RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion

Zhe Li, Cheng Chi, Boan Zhu +11

Humans learn locomotion through visual observation, interpreting visual content first before imitating actions. However, state-of-the-art humanoid locomotion systems rely on either…

cs.CV2025

ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better

Yuan Zhang, Ming Lu, Junwen Pan +4

Recent advances in multimodal reasoning models have demonstrated impressive capabilities across text and vision. However, even leading models exhibit redundant self-reflection when…

cs.RO2025

From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance

Zhe Li, Cheng Chi, Yangyang Wei +7

Natural language offers a natural interface for humanoid robots, but existing language-guided humanoid locomotion pipelines remain cumbersome and untrustworthy. They typically deco…

cs.CV2025

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Yuan Zhang, Chun-Kai Fan, Junpeng Ma +8

In vision-language models (VLMs), visual tokens usually bear a significant amount of computational overhead despite sparsity of information in them when compared to text tokens. To…