collaborators

5 papers

cs.CV2026

Occlusion-Robust Multi-Object Decoupling for Physics-Based Robotic Interaction

Xin Dong, Lihan Zhang, Tianru Dai +2

We propose a mask-free method for lossless multi-object 3D reconstruction from sparse and occluded real-world views, enabling physically plausible robotic interaction via Material…

cs.CV2026

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection

Xin Dong, Wenjia Geng, Wenfeng Deng +1

Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given…

cs.CV2026

CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction

Xin Dong, Weijian Deng, Lihan Zhang +3

Reconstructing interaction-ready 3D worlds is essential for physical simulation, virtual reality, robotics, and autonomous driving. However, existing methods mainly optimize static…

cs.CV2026

Boosting Zero-Shot 3D Style Transfer with 2D Pre-trained Priors

Xin Dong, Yunzhi Teng, Wenfeng Deng +1

In this work, we focus on zero-shot 3D style transfer that can generate multi-view consistent stylized views of the 3D scene given an arbitrary style image. We primarily tackle the…

cs.CV2025

Localization-Aware Multi-Scale Representation Learning for Repetitive Action Counting

Sujia Wang, Xiangwei Shen, Yansong Tang +3

Repetitive action counting (RAC) aims to estimate the number of class-agnostic action occurrences in a video without exemplars. Most current RAC methods rely on a raw frame-to-fram…