4 papers
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
Jixuan He, Xueting Li, Chieh Hubert Lin +1
Vision-Language Models (VLMs) exhibit emerging spatial reasoning capabilities, yet they remain unreliable on tasks requiring precise spatial understanding, such as viewpoint reason…
Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video
Jixuan He, Chieh Hubert Lin, Lu Qi +1
Creating deformable 3D content has gained increasing attention with the rise of text-to-image and image-to-video generative models. While these models provide rich semantic priors…
Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion
Jixuan He, Wanhua Li, Ye Liu +3
As a common image editing operation, image composition involves integrating foreground objects into background scenes. In this paper, we expand the application of the concept of Af…
-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding
Ye Liu, Jixuan He, Wanhua Li +4
Video temporal grounding (VTG) is a fine-grained video understanding problem that aims to ground relevant clips in untrimmed videos given natural language queries. Most existing VT…