collaborators

10 papers

cs.CV2026

Direct 3D-Aware Object Insertion via Decomposed Visual Proxies

Jingbo Gong, Yikai Wang, Yushi Lan +6

Object insertion aims to seamlessly composite a reference object into a specified region of a background image. Recent diffusion-based methods achieve high visual quality but formu…

cs.CV2026

4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere

Yihang Luo, Shangchen Zhou, Yushi Lan +2

We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing approaches that typically decouple motion from geometry or produce lim…

cs.CV2025

Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence

Shuai Yang, Junxin Lin, Yifan Zhou +2

The remarkable success in text-to-image diffusion models has motivated extensive investigation of their potential for video applications. Zero-shot techniques aim to adapt image di…

cs.RO2025

DHAGrasp: Synthesizing Affordance-Aware Dual-Hand Grasps with Text Instructions

Quanzhou Li, Zhonghua Wu, Jingbo Wang +2

Learning to generate dual-hand grasps that respect object semantics is essential for robust hand-object interaction but remains largely underexplored due to dataset scarcity. Exist…

cs.CV2025

STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer

Yushi Lan, Yihang Luo, Fangzhou Hong +7

We present STream3R, a novel approach to 3D reconstruction that reformulates pointmap prediction as a decoder-only Transformer problem. Existing state-of-the-art methods for multi-…

cs.CV2025

GausSim: Foreseeing Reality by Gaussian Simulator for Elastic Objects

Yidi Shao, Mu Huang, Chen Change Loy +1

We introduce GausSim, a novel neural network-based simulator designed to capture the dynamic behaviors of real-world elastic objects represented through Gaussian kernels. We levera…