11 papers
DynActiveGS: Active Gaussian Splatting for Dynamic Scene Reconstruction
Hongbo Duan, Pengting Luo, Chengzhi Zhao +3
We present DynActiveGS, a dynamic-aware active reconstruction framework based on 3D Gaussian Splatting (3DGS) for autonomous exploration in dynamic environments. The framework incr…
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
Jiacheng Hou, Yining Sun, Ruochong Jin +4
Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual i…
VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks
Yining Sun, Haoyu Kang, Jiajun Wu +7
Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as…
Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs
Haochen Han, Jue Wang, Alex Jinpeng Wang +1
Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in various Remote Sensing (RS) tasks. However, their ability to comprehend negation remains underexplo…
Pulse: Training Acceleration for Large Diffusion Models with Automatic Pipeline Parallelism
Boran Sun, Guoyong Jiang, Lin Zhang +8
Diffusion models are now a dominant approach for high-fidelity image and video generation, yet scaling their training across GPU clusters remains challenging. Unlike transformer-on…
PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation
Lingxuan Wu, Zijian Zhu, Lizhong Wang +5
Diffusion policies have achieved remarkable success in robotic manipulation, yet they often fail to satisfy strict physical constraints required for safe deployment. Existing appro…