9 papers
VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
Jiazi Wang, Nonghai Zhang, Qiushi Xie +5
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cult…
Representation Forcing for Bottleneck-Free Unified Multimodal Models
Yuqing Wang, Zhijie Lin, Ceyuan Yang +10
Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation…
SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction
Yiran Wang, Zeyu Zhang, Yuanming Li +2
High-quality 4D head avatars from one or a few source portraits are central to telepresence, AR/VR, and digital-human interaction. 3D Gaussian Splatting (3DGS) has emerged as the d…
A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-Rational Cognition, and Governance of AI-Generated Content
Yang Zhao, Yingshuo Li, Zeyu Zhang
Large language models are increasingly used for knowledge acquisition, code generation, academic writing, and agent-based automation. In these settings, users may obtain highly str…
PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps
Junlin Long, Zeyu Zhang, Xu Deng +5
Embodied visual navigation, where an agent perceives a complex environment and acts to reach a goal from raw sensory input, underpins a wide range of applications such as household…
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
Wei Wu, Ziyang Xu, Zeyu Zhang +2
Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal media, and interactive delivery.…