4 papers
Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
Yiwen Tang, Zoey Guo, Kaixin Zhu +11
Reinforcement learning (RL), earlier proven to be effective in large language and multi-modal models, has been successfully extended to enhance 2D image generation recently. Howeve…
FastUMI-100K: Advancing Data-driven Robotic Manipulation with a Large-scale UMI-style Dataset
Kehui Liu, Zhongjie Jia, Yang Li +14
Data-driven robotic manipulation learning depends on large-scale, high-quality expert demonstration datasets. However, existing datasets, which primarily rely on human teleoperated…
Trajectory Conditioned Cross-embodiment Skill Transfer
YuHang Tang, Yixuan Lou, Pengfei Han +4
Learning manipulation skills from human demonstration videos presents a promising yet challenging problem, primarily due to the significant embodiment gap between human body and ro…
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
Hailong Ning, Siying Wang, Tao Lei +5
Remote Sensing Image-Text Retrieval (RSITR) plays a critical role in geographic information interpretation, disaster monitoring, and urban planning by establishing semantic associa…