From the 1 of 5 linked papers with an AI index.
5 papers
Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery
Xinliang Wang, Yijie Kang, Zhenyu Wu +1
Sat2RealCity is a framework that generates 3D urban models from satellite images by grounding object-level 3D generative priors to real-world geographic locations and allowing cont…
DockAnywhere: Data-Efficient Visuomotor Policy Learning for Mobile Manipulation via Novel Demonstration Generation
Ziyu Shan, Yuheng Zhou, Gaoyuan Wu +3
Mobile manipulation is a fundamental capability that enables robots to interact in expansive environments such as homes and factories. Most existing approaches follow a two-stage p…
ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models
Xinliang Wang, Yifeng Shi, Zhenyu Wu
3D Gaussian Splatting (3DGS) delivers high-fidelity real-time rendering but suffers from geometric and photometric degradations under sparse-view constraints. Current generative re…
VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
Shiji Zhao, Shukun Xiong, Yao Huang +7
Multimodal Large Language Models (MLLMs) are widely used in various fields due to their powerful cross-modal comprehension and generation capabilities. However, more modalities bri…
The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation
Shouwei Ruan, Zhenyu Wu, Yao Huang +5
Content safety is a fundamental challenge for text-to-image (T2I) models, yet prevailing methods enforce a debilitating trade-off between safety and generation quality. We argue th…