8 papers
Occlusion-Robust Multi-Object Decoupling for Physics-Based Robotic Interaction
Xin Dong, Lihan Zhang, Tianru Dai +2
We propose a mask-free method for lossless multi-object 3D reconstruction from sparse and occluded real-world views, enabling physically plausible robotic interaction via Material…
CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection
Xin Dong, Wenjia Geng, Wenfeng Deng +1
Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given…
CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction
Xin Dong, Weijian Deng, Lihan Zhang +3
Reconstructing interaction-ready 3D worlds is essential for physical simulation, virtual reality, robotics, and autonomous driving. However, existing methods mainly optimize static…
Boosting Zero-Shot 3D Style Transfer with 2D Pre-trained Priors
Xin Dong, Yunzhi Teng, Wenfeng Deng +1
In this work, we focus on zero-shot 3D style transfer that can generate multi-view consistent stylized views of the 3D scene given an arbitrary style image. We primarily tackle the…
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
Shilin Ma, Chubin Zhang, Changyuan Wang +6
Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most…
Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
Xiangwei Shen, Zhimin Li, Zhantao Yang +6
Recent studies have demonstrated the effectiveness of directly aligning diffusion models with human preferences using differentiable reward. However, they exhibit two primary chall…