collaborators

8 papers

cs.CV2026

Occlusion-Robust Multi-Object Decoupling for Physics-Based Robotic Interaction

Xin Dong, Lihan Zhang, Tianru Dai +2

We propose a mask-free method for lossless multi-object 3D reconstruction from sparse and occluded real-world views, enabling physically plausible robotic interaction via Material…

cs.CV2026

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection

Xin Dong, Wenjia Geng, Wenfeng Deng +1

Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given…

cs.CV2026

CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction

Xin Dong, Weijian Deng, Lihan Zhang +3

Reconstructing interaction-ready 3D worlds is essential for physical simulation, virtual reality, robotics, and autonomous driving. However, existing methods mainly optimize static…

cs.CV2026

Boosting Zero-Shot 3D Style Transfer with 2D Pre-trained Priors

Xin Dong, Yunzhi Teng, Wenfeng Deng +1

In this work, we focus on zero-shot 3D style transfer that can generate multi-view consistent stylized views of the 3D scene given an arbitrary style image. We primarily tackle the…

cs.CV2026

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

Shilin Ma, Chubin Zhang, Changyuan Wang +6

Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most…

cs.AI2025

Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference

Xiangwei Shen, Zhimin Li, Zhantao Yang +6

Recent studies have demonstrated the effectiveness of directly aligning diffusion models with human preferences using differentiable reward. However, they exhibit two primary chall…