8 papers · 1 filter
Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking
Zhengbo Zhang, Zhigang Tu, Junsong Yuan +2
Unsupervised visual object tracking is a challenging task that requires following arbitrary targets in videos without training on ground-truth annotations. Despite considerable pro…
VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation
Jiayi Yuan, Haobo Jiang, De Wen Soh +1
This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-independent training-free approaches,…
OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects
Mark He Huang, Lin Geng Foo, Christian Theobalt +2
Free-moving object reconstruction from monocular video remains challenging, particularly without reliable pose or depth cues and under arbitrary object motion. We introduce OnlineS…
MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm
Ziyan Guo, Zeyu Hu, De Wen Soh +1
Human motion generation and editing are key components of computer vision. However, current approaches in this field tend to offer isolated solutions tailored to specific tasks, wh…
TSTMotion: Training-free Scene-aware Text-to-motion Generation
Ziyan Guo, Haoxuan Qu, Hossein Rahmani +4
Text-to-motion generation has recently garnered significant research interest, primarily focusing on generating human motion sequences in blank backgrounds. However, human motions…
GaussianBlock: Building Part-Aware Compositional and Editable 3D Scene by Primitives and Gaussians
Shuyi Jiang, Qihao Zhao, Hossein Rahmani +3
Recently, with the development of Neural Radiance Fields and Gaussian Splatting, 3D reconstruction techniques have achieved remarkably high fidelity. However, the latent representa…