7 papers
Coherent and Multi-modality Image Inpainting via Latent Space Optimization
Lingzhi Pan, Tong Zhang, Bingyuan Chen +4
With the advancements in denoising diffusion probabilistic models (DDPMs), image inpainting has significantly evolved from merely filling information based on nearby regions to gen…
Demystifying Singular Defects in Large Language Models
Haoqi Wang, Tong Zhang, Mathieu Salzmann
Large transformer models are known to produce high-norm tokens. In vision transformers (ViTs), such tokens have been mathematically modeled through the singular vectors of the line…
Adaptive Multi-step Refinement Network for Robust Point Cloud Registration
Zhi Chen, Yufan Ren, Tong Zhang +4
Point Cloud Registration (PCR) estimates the relative rigid transformation between two point clouds of the same scene. Despite significant progress with learning-based approaches,…
DVMNet++: Rethinking Relative Pose Estimation for Unseen Objects
Chen Zhao, Tong Zhang, Zheng Dang +1
Determining the relative pose of a previously unseen object between two images is pivotal to the success of generalizable object pose estimation. Existing approaches typically pred…
Self-Ensembling Gaussian Splatting for Few-Shot Novel View Synthesis
Chen Zhao, Xuan Wang, Tong Zhang +2
3D Gaussian Splatting (3DGS) has demonstrated remarkable effectiveness in novel view synthesis (NVS). However, 3DGS tends to overfit when trained with sparse views, limiting its ge…
Unlocking Comics: The AI4VA Dataset for Visual Understanding
Peter Grönquist, Deblina Bhattacharjee, Bahar Aydemir +4
In the evolving landscape of deep learning, there is a pressing need for more comprehensive datasets capable of training models across multiple modalities. Concurrently, in digital…