papers

Publications (57)

cs.CV2024

RIAV-MVS: Recurrent-Indexing an Asymmetric Volume for Multi-View Stereo

Changjiang Cai, Pan Ji, Qingan Yan +1

This paper presents a learning-based method for multi-view depth estimation from posed images. Our core idea is a "learning-to-optimize" paradigm that iteratively indexes a plane-s…

cs.CV2026

Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation

Hanlin Chen, Jiaxin Wei, Xibin Song +5

Frame-wise action-controlled image-to-video generation is a promising paradigm for interactive world simulation, where each control signal should elicit an immediate visual respons…

cs.CV2025

MARS: Mesh AutoRegressive Model for 3D Shape Detailization

Jingnan Gao, Weizhe Liu, Weixuan Sun +8

State-of-the-art methods for mesh detailization predominantly utilize Generative Adversarial Networks (GANs) to generate detailed meshes from coarse ones. These methods typically l…

cs.CV2025

CDI3D: Cross-guided Dense-view Interpolation for 3D Reconstruction

Zhiyuan Wu, Xibin Song, Senbo Wang +8

3D object reconstruction from single-view image is a fundamental task in computer vision with wide-ranging applications. Recent advancements in Large Reconstruction Models (LRMs) h…

cs.RO2022

CNN-Augmented Visual-Inertial SLAM with Planar Constraints

Pan Ji, Yuan Tian, Qingan Yan +2

We present a robust visual-inertial SLAM system that combines the benefits of Convolutional Neural Networks (CNNs) and planar constraints. Our system leverages a CNN to predict the…

cs.CV2021

MonoIndoor: Towards Good Practice of Self-Supervised Monocular Depth Estimation for Indoor Environments

Pan Ji, Runze Li, Bir Bhanu +1

Self-supervised depth estimation for indoor environments is more challenging than its outdoor counterpart in at least the following two aspects: (i) the depth range of indoor seque…