11 papers
Keep the Future, Drop the Rollout: RIFT for World Action Models
Chushan Zhang, Jinguang Tong, Xuesong Li +2
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evo…
Label Shift Aware Adaptation for Online Zero-shot Learning with Contrastive Language-Image Pre-Training (CLIP)
Pengxiao Han, Changkun Ye, Yanshuo Wang +5
Vision-language models like Contrastive Language-Image Pre-Training (CLIP) have been extensively studied in data-scarce scenarios. A particularly challenging and realistic task in…
Structural Energy Guidance for View-Consistent Text-to-3D Generation
Qing Zhang, Jinguang Tong, Jing Zhang +2
Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work identifies viewpoint bias in 2D…
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
Chushan Zhang, Ruihan Lu, Jinguang Tong +3
Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet robot actions cause contact,…
MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
Jinguang Tong, Jinbo Wu, Kaisiyuan Wang +12
Human-Object Interaction (HOI) video reenactment aims to transfer the interaction dynamics of a source video to a novel target object while preserving realistic hand-object coordin…
Adaptive and Balanced Re-initialization for Long-timescale Continual Test-time Domain Adaptation
Yanshuo Wang, Jinguang Tong, Jun Lan +5
Continual test-time domain adaptation (CTTA) aims to adjust models so that they can perform well over time across non-stationary environments. While previous methods have made cons…