From the 1 of 7 linked papers with an AI index.
7 papers
MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
Ziren Gong, Xiaohan Li, Fabio Tosi +4
The paper introduces MAGiSt3R, a multi-agent framework that reconstructs 3D scenes and tracks camera pose from monocular RGB videos in near real-time using feed-forward models and…
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
Yusheng Dai, Zehua Chen, Yuxuan Jiang +4
Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet face…
MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE
Ruijie Zhu, Jiahao Lu, Wenbo Hu +4
We present MotionCrafter, a framework that leverages video generators to jointly reconstruct 4D geometry and estimate dense motion from a monocular video. The key idea is a joint r…
Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
Ziren Gong, Xiaohan Li, Fabio Tosi +4
We present Ov3R, a novel framework for open-vocabulary semantic 3D reconstruction from RGB video streams, designed to advance Spatial AI. The system features two key components: CL…
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
Chenhui Gou, Zilong Chen, Zeyu Wang +10
This paper studies Visual Question-Visual Answering (VQ-VA): generating an image, rather than text, in response to a visual question -- an ability that has recently emerged in prop…
Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis
Hongyu Sun, Qiuhong Ke, Ming Cheng +4
This paper proposes a general solution to enable point cloud recognition models to handle distribution shifts at test time. Unlike prior methods, which rely heavily on training dat…