12 papers
PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing
Ruihang Xu, Dewei Zhou, Xiaolong Shen +2
Achieving physically accurate object manipulation in image editing is essential for its potential applications in interactive world models. However, existing visual generative mode…
Echoes of ownership: Adversarial-guided dual injection for copyright protection in MLLMs
Chengwei Xia, Fan Ma, Ruijie Quan +3
With the rapid deployment of multimodal large language models (MLLMs), disputes regarding model ownership have become increasingly frequent, raising significant concerns about inte…
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
Ruihang Xu, Dewei Zhou, Fan Ma +1
Multi-instance image generation (MIG) remains a significant challenge for modern diffusion models due to key limitations in achieving precise control over object layout and preserv…
Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
Yixuan Han, Fan Ma, Ruijie Quan +1
Test-Time Scaling (TTS) enhances the reasoning ability of large language models (LLMs) by allocating additional computation during inference. However, existing approaches primarily…
Adversarial-Guided Diffusion for Multimodal LLM Attacks
Chengwei Xia, Fan Ma, Ruijie Quan +2
This paper addresses the challenge of generating adversarial image using a diffusion model to deceive multimodal large language models (MLLMs) into generating the targeted response…
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
Dewei Zhou, You Li, Fan Ma +2
We introduce the Multi-Instance Generation (MIG) task, which focuses on generating multiple instances within a single image, each accurately placed at predefined positions with att…