papers

Publications (15)

cs.RO2023

MOMA-Force: Visual-Force Imitation for Real-World Mobile Manipulation

Taozheng Yang, Ya Jing, Hongtao Wu +5

In this paper, we present a novel method for mobile manipulators to perform multiple contact-rich manipulation tasks. While learning-based methods have the potential to generate ac…

cs.CL2024

Knowledge Boundary and Persona Dynamic Shape A Better Social Media Agent

Junkai Zhou, Liang Pang, Ya Jing +3

Constructing personalized and anthropomorphic agents holds significant importance in the simulation of social networks. However, there are still two key problems in existing works:…

cs.RO2023

Learning to Explore Informative Trajectories and Samples for Embodied Perception

Ya Jing, Tao Kong

We are witnessing significant progress on perception models, specifically those trained on large-scale internet images. However, efficiently generalizing these perception models to…

cs.CV2021

Locate then Segment: A Strong Pipeline for Referring Image Segmentation

Ya Jing, Tao Kong, Wei Wang +3

Referring image segmentation aims to segment the objects referred by a natural language expression. Previous methods usually focus on designing an implicit and recurrent feature in…

cs.CV2026

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning

Liangyu Fu, Junbo Wang, Yuke Li +3

Text-only training is a popular paradigm in zero-shot video captioning, where the video distribution is not available to the model during training, leading to a cross-modal gap bet…

cs.RO2023

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Hongtao Wu, Ya Jing, Chilam Cheang +6

Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of th…