papers

Publications (13)

cs.CV2023

Implicit Autoencoder for Point-Cloud Self-Supervised Representation Learning

Siming Yan, Zhenpei Yang, Haoxiang Li +5

This paper advocates the use of implicit surface representation in autoencoder-based self-supervised 3D representation learning. The most popular and accessible 3D representation,…

cs.CV2025

ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling

Siming Yan, Min Bai, Weifeng Chen +3

By combining natural language understanding, generation capabilities, and breadth of knowledge of large language models with image perception, recent large vision language models (…

cs.CV2020

Extreme Relative Pose Network under Hybrid Representations

Zhenpei Yang, Siming Yan, Qixing Huang

In this paper, we introduce a novel RGB-D based relative pose estimation approach that is suitable for small-overlapping or non-overlapping scans and can output multiple relative p…

cs.CV2024

Multi-View Representation is What You Need for Point-Cloud Pre-Training

Siming Yan, Chen Song, Youkang Kong +1

A promising direction for pre-training 3D point clouds is to leverage the massive amount of data in 2D, whereas the domain gap between 2D and 3D creates a fundamental challenge. Th…

cs.CV2024

3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud Pretraining

Siming Yan, Yuqi Yang, Yuxiao Guo +5

Masked autoencoders (MAE) have recently been introduced to 3D self-supervised pretraining for point clouds due to their great success in NLP and computer vision. Unlike MAEs used i…

cs.CV2021

Scene Synthesis via Uncertainty-Driven Attribute Synchronization

Haitao Yang, Zaiwei Zhang, Siming Yan +5

Developing deep neural networks to generate 3D scenes is a fundamental problem in neural synthesis with immediate applications in architectural CAD, computer graphics, as well as i…

cs.CV2018

Calcium Removal From Cardiac CT Images Using Deep Convolutional Neural Network

Siming Yan, Feng Shi, Yuhua Chen +5

Coronary calcium causes beam hardening and blooming artifacts on cardiac computed tomography angiography (CTA) images, which lead to overestimation of lumen stenosis and reduction…

cs.CV2021

HPNet: Deep Primitive Segmentation Using Hybrid Representations

Siming Yan, Zhenpei Yang, Chongyang Ma +3

This paper introduces HPNet, a novel deep-learning approach for segmenting a 3D shape represented as a point cloud into primitive patches. The key to deep primitive segmentation is…

cs.CV2026

DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding

Dong Zhuo, Wenzhao Zheng, Sicheng Zuo +4

With the growing adoption of vision-language-action models and world models in autonomous driving systems, scalable image tokenization becomes crucial as the interface for the visu…

cs.CV2025

DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning

Zhe Liu, Runhui Huang, Rui Yang +6

Although multi-modal large language models (MLLMs) have shown strong capabilities across diverse domains, their application in generating fine-grained 3D perception and prediction…

cs.CV2025

Representation Learning for Point Cloud Understanding

Siming Yan

With the rapid advancement of technology, 3D data acquisition and utilization have become increasingly prevalent across various fields, including computer vision, robotics, and geo…

q-bio.NC2019

Recurrent Feedback Improves Feedforward Representations in Deep Neural Networks

Siming Yan, Xuyang Fang, Bowen Xiao +3

The abundant recurrent horizontal and feedback connections in the primate visual cortex are thought to play an important role in bringing global and semantic contextual information…

cs.CV2025

Wavelet-based Decoupling Framework for low-light Stereo Image Enhancement

Shuangli Du, Siming Yan, Zhenghao Shi +2

Low-light images suffer from complex degradation, and existing enhancement methods often encode all degradation factors within a single latent space. This leads to highly entangled…