6 papers
Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation
Jingyi Lu, Kai Han
Monocular-to-stereo conversion synthesizes stereoscopic content from 2D videos for immersive 3D experiences. In modern Depth-Image-Based Rendering (DIBR) approaches, stereo inpaint…
Semantic Correspondence: Unified Benchmarking and a Strong Baseline
Kaiyan Zhang, Xinghui Li, Jingyi Lu +1
Establishing semantic correspondence is a challenging task in computer vision, aiming to match keypoints with the same semantic information across different images. Benefiting from…
SCoPE: Sightline-Coordinate Positional Encoding for Video Diffusion Transformers
Minghao Yin, Jiahao Lu, Wenbo Hu +3
Video diffusion transformers address their tokens by position on the pixel-time grid: an address in the tensor, not in the world. The address we would want, the world point a token…
Scene-Level Heterogeneous Physics Simulation with 3D Gaussian Splats
Xiaoyang Liu, Shangzhe Wu, Kai Han
3D Gaussian Splatting (3DGS) has achieved state-of-the-art photorealistic rendering, but the representation gap prevents these assets from being physically interactive. Production-…
VAGS: Velocity Adaptive Guidance Scale for Image Editing and Generation
Yan Luo, Ahmadou Aidara, Jingyi Lu +3
Classifier-free guidance (CFG) is the primary control over how strongly text semantics move a flow-based sampler, yet standard practice holds its scale fixed across the entire ODE…
Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping
Jingyi Lu, Kai Han
Drag-based image editing has emerged as a powerful paradigm for intuitive image manipulation. However, existing approaches predominantly rely on manipulating the latent space of ge…