7 papers
PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling
Zhijie Zheng, Xinhao Xiang, Jiawei Zhang
Reconstructing 3D scenes from a single image is a fundamental challenge in computer vision, with broad applications in virtual reality, robotics, and content creation. Recent metho…
TTSA3R: Training-Free Temporal-Spatial Adaptive Persistent State for Streaming 3D Reconstruction
Zhijie Zheng, Xinhao Xiang, Jiawei Zhang
Streaming recurrent models enable efficient 3D reconstruction by maintaining persistent state representations. However, they suffer from catastrophic forgetting over long sequences…
A Survey of AI-Generated Video Evaluation
Xiao Liu, Xinhao Xiang, Zizhong Li +6
The growing capabilities of AI in generating video content have brought forward significant challenges in effectively evaluating these videos. Unlike static images or text, video c…
OccFace: Unified Occlusion-Aware Facial Landmark Detection with Per-Point Visibility
Xinhao Xiang, Zhengxin Li, Saurav Dhakad +3
Accurate facial landmark detection under occlusion remains challenging, especially for human-like faces with large appearance variation and rotation-driven self-occlusion. Existing…
Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework
Xinhao Xiang, Abhijeet Rastogi, Jiawei Zhang
Recent text-to-video models have enabled the generation of high-resolution driving scenes from natural language prompts. These AI-generated driving videos (AIGVs) offer a low-cost,…
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
Xinhao Xiang, Kuan-Chuan Peng, Suhas Lohit +2
3D object detection plays a crucial role in autonomous systems, yet existing methods are limited by closed-set assumptions and struggle to recognize novel objects and their attribu…