3 citations · 3 across the 5 of their papers we have counts for
4 papers · 1 filter
VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
Jiazi Wang, Nonghai Zhang, Qiushi Xie +5
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cult…
MotionVLA: Vision-Language-Action Model for Humanoid Motion
Nonghai Zhang, Siyu Zhai, Yanjun Li +5
Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-frequency physical dynamics. However, many existing methods toke…
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
Nonghai Zhang, Zeyu Zhang, Jiazi Wang +2
Vision-Language Models (VLMs) have achieved significant progress in multimodal understanding tasks, demonstrating strong capabilities particularly in general tasks such as image ca…
Text-to-Image Synthesis: A Decade Survey
Nonghai Zhang, Hao Tang
When humans read a specific text, they often visualize the corresponding images, and we hope that computers can do the same. Text-to-image synthesis (T2I), which focuses on generat…