Publications (15)
MOMA-Force: Visual-Force Imitation for Real-World Mobile Manipulation
Taozheng Yang, Ya Jing, Hongtao Wu +5
In this paper, we present a novel method for mobile manipulators to perform multiple contact-rich manipulation tasks. While learning-based methods have the potential to generate ac…
Knowledge Boundary and Persona Dynamic Shape A Better Social Media Agent
Junkai Zhou, Liang Pang, Ya Jing +3
Constructing personalized and anthropomorphic agents holds significant importance in the simulation of social networks. However, there are still two key problems in existing works:…
Learning to Explore Informative Trajectories and Samples for Embodied Perception
Ya Jing, Tao Kong
We are witnessing significant progress on perception models, specifically those trained on large-scale internet images. However, efficiently generalizing these perception models to…
Locate then Segment: A Strong Pipeline for Referring Image Segmentation
Ya Jing, Tao Kong, Wei Wang +3
Referring image segmentation aims to segment the objects referred by a natural language expression. Previous methods usually focus on designing an implicit and recurrent feature in…
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning
Liangyu Fu, Junbo Wang, Yuke Li +3
Text-only training is a popular paradigm in zero-shot video captioning, where the video distribution is not available to the model during training, leading to a cross-modal gap bet…
Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
Hongtao Wu, Ya Jing, Chilam Cheang +6
Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of th…