5 papers
Occlusion-Robust Multi-Object Decoupling for Physics-Based Robotic Interaction
Xin Dong, Lihan Zhang, Tianru Dai +2
We propose a mask-free method for lossless multi-object 3D reconstruction from sparse and occluded real-world views, enabling physically plausible robotic interaction via Material…
STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
Haipeng Luo, Qingfeng Sun, Songli Wu +4
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from poli…
CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection
Xin Dong, Wenjia Geng, Wenfeng Deng +1
Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given…
CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction
Xin Dong, Weijian Deng, Lihan Zhang +3
Reconstructing interaction-ready 3D worlds is essential for physical simulation, virtual reality, robotics, and autonomous driving. However, existing methods mainly optimize static…
Boosting Zero-Shot 3D Style Transfer with 2D Pre-trained Priors
Xin Dong, Yunzhi Teng, Wenfeng Deng +1
In this work, we focus on zero-shot 3D style transfer that can generate multi-view consistent stylized views of the 3D scene given an arbitrary style image. We primarily tackle the…