5 papers
TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition
Minghao Zhu, Xiao Lin, Mengxian Hu +5
Adapting CLIP for open-vocabulary video recognition necessitates a delicate balance between newly acquired video knowledge and the pretrained generalization. While existing studies…
CLASH: Collaborative Large-Small Hierarchical Framework for Continuous Vision-and-Language Navigation
Liuyi Wang, Zongtao He, Jinlong Li +6
Vision-and-Language Navigation (VLN) requires robots to follow natural language instructions and navigate complex environments without prior maps. While recent vision-language larg…
Realizing Text-Driven Motion Generation on NAO Robot: A Reinforcement Learning-Optimized Control Pipeline
Zihan Xu, Mengxian Hu, Kaiyan Xiao +3
Human motion retargeting for humanoid robots, transferring human motion data to robots for imitation, presents significant challenges but offers considerable potential for real-wor…
Efficient Text-driven Motion Generation via Latent Consistency Training
Mengxian Hu, Minghao Zhu, Xun Zhou +4
Text-driven human motion generation based on diffusion strategies establishes a reliable foundation for multimodal applications in human-computer interactions. However, existing ad…
MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer
Minghao Zhu, Zhengpu Wang, Mengxian Hu +5
Transferring visual-language knowledge from large-scale foundation models for video recognition has proved to be effective. To bridge the domain gap, additional parametric modules…