11 papers
Keep the Future, Drop the Rollout: RIFT for World Action Models
Chushan Zhang, Jinguang Tong, Xuesong Li +2
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evo…
Maintain Plasticity in Long-timescale Continual Test-time Adaptation
Yanshuo Wang, Xuesong Li, Jinguang Tong +5
Continual test-time domain adaptation (CTTA) aims to adjust pre-trained source models to perform well over time across non-stationary target environments. While previous methods ha…
Label Shift Aware Adaptation for Online Zero-shot Learning with Contrastive Language-Image Pre-Training (CLIP)
Pengxiao Han, Changkun Ye, Yanshuo Wang +5
Vision-language models like Contrastive Language-Image Pre-Training (CLIP) have been extensively studied in data-scarce scenarios. A particularly challenging and realistic task in…
Structural Energy Guidance for View-Consistent Text-to-3D Generation
Qing Zhang, Jinguang Tong, Jing Zhang +2
Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work identifies viewpoint bias in 2D…
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
Chushan Zhang, Ruihan Lu, Jinguang Tong +3
Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet robot actions cause contact,…
MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
Jinguang Tong, Jinbo Wu, Kaisiyuan Wang +12
Human-Object Interaction (HOI) video reenactment aims to transfer the interaction dynamics of a source video to a novel target object while preserving realistic hand-object coordin…