185 citations · 282 across the 11 of their papers we have counts for
9 papers
X-3D: Explicit 3D Structure Modeling for Point Cloud Recognition
Shuofeng Sun, Yongming Rao, Jiwen Lu +1
Numerous prior studies predominantly emphasize constructing relation vectors for individual neighborhood points and generating dynamic kernels for each vector and embedding these i…
Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Zuyan Liu, Yuhao Dong, Yongming Rao +2
In the realm of vision-language understanding, the proficiency of models in interpreting and reasoning over visual content has become a cornerstone for numerous applications. Howev…
TCOVIS: Temporally Consistent Online Video Instance Segmentation
Junlong Li, Bingyao Yu, Yongming Rao +2
In recent years, significant progress has been made in video instance segmentation (VIS), with many offline and online methods achieving state-of-the-art performance. While offline…
Unleashing Text-to-Image Diffusion Models for Visual Perception
Wenliang Zhao, Yongming Rao, Zuyan Liu +3
Diffusion models (DMs) have become the new trend of generative models and have demonstrated a powerful ability of conditional synthesis. Among those, text-to-image diffusion models…
AdaPoinTr: Diverse Point Cloud Completion with Adaptive Geometry-Aware Transformers
Xumin Yu, Yongming Rao, Ziyi Wang +2
In this paper, we present a new method that reformulates point cloud completion as a set-to-set translation problem and design a new model, called PoinTr, which adopts a Transforme…
P2P: Tuning Pre-trained Image Models for Point Cloud Analysis with Point-to-Pixel Prompting
Ziyi Wang, Xumin Yu, Yongming Rao +2
Nowadays, pre-training big models on large-scale datasets has become a crucial topic in deep learning. The pre-trained models with high representation ability and transferability a…