70 citations · 127 across the 8 of their papers we have counts for
9 papers
MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion
Shitao Tang, Fuyang Zhang, Jiacheng Chen +2
This paper introduces MVDiffusion, a simple yet effective method for generating consistent multi-view images from text prompts given pixel-to-pixel correspondences (e.g., perspecti…
NeuMap: Neural Coordinate Mapping by Auto-Transdecoder for Camera Localization
Shitao Tang, Sicong Tang, Andrea Tagliasacchi +2
This paper presents an end-to-end neural mapping method for camera localization, dubbed NeuMap, encoding a whole scene into a grid of latent codes, with which a Transformer-based a…
RenderNet: Visual Relocalization Using Virtual Viewpoints in Large-Scale Indoor Environments
Jiahui Zhang, Shitao Tang, Kejie Qiu +6
Visual relocalization has been a widely discussed problem in 3D vision: given a pre-constructed 3D visual map, the 6 DoF (Degrees-of-Freedom) pose of a query image is estimated. Re…
QuadTree Attention for Vision Transformers
Shitao Tang, Jiahui Zhang, Siyu Zhu +1
Transformers have been successful in many vision tasks, thanks to their capability of capturing long-range dependency. However, their quadratic computational complexity poses a maj…
Learning Camera Localization via Dense Scene Matching
Shitao Tang, Chengzhou Tang, Rui Huang +2
Camera localization aims to estimate 6 DoF camera poses from RGB images. Traditional methods detect and match interest points between a query image and a pre-built 3D model. Recent…
Channel Equilibrium Networks for Learning Deep Representation
Wenqi Shao, Shitao Tang, Xingang Pan +3
Convolutional Neural Networks (CNNs) are typically constructed by stacking multiple building blocks, each of which contains a normalization layer such as batch normalization (BN) a…