3 citations · 6 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 3 cited
LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Ruyi Xu, Yuan Yao, Zonghao Guo +7
Visual encoding constitutes the basis of large multimodal models (LMMs) in understanding the visual world. Conventional LMMs process images in fixed sizes and limited resolutions,…
cs.CV2024★ 1 cited
ControlCap: Controllable Region-level Captioning
Yuzhong Zhao, Yue Liu, Zonghao Guo +4
Region-level captioning is challenged by the caption degeneration issue, which refers to that pre-trained multimodal models tend to predict the most frequent captions but miss the…
cs.CV2022★ 2 cited
Bidirectional Feature Globalization for Few-shot Semantic Segmentation of 3D Point Cloud Scenes
Yongqiang Mao, Zonghao Guo, Xiaonan Lu +2
Few-shot segmentation of point cloud remains a challenging task, as there is no effective way to convert local point cloud information to global representation, which hinders the g…