5 papers
VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
Jiazi Wang, Nonghai Zhang, Qiushi Xie +5
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cult…
DualGeo: A Dual-View Framework for Worldwide Image Geo-localization
Junchao Cui, Wenqi Shi, Shaoyong Du +4
Worldwide image geo-localization aims to infer the geographic location of an image captured anywhere on Earth, spanning street, city, regional, national, and continental scales. Ex…
Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs
Chongyu Wang, Ting Huang, Chunyu Sun +3
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in 2D visual tasks but still struggle to understand physical space in real-world visual streams. Recently…
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
Nonghai Zhang, Zeyu Zhang, Jiazi Wang +2
Vision-Language Models (VLMs) have achieved significant progress in multimodal understanding tasks, demonstrating strong capabilities particularly in general tasks such as image ca…
Text-to-Image Synthesis: A Decade Survey
Nonghai Zhang, Hao Tang
When humans read a specific text, they often visualize the corresponding images, and we hope that computers can do the same. Text-to-image synthesis (T2I), which focuses on generat…