3 citations · 5 across the 25 of their papers we have counts for
9 papers · 1 filter
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
Chuyan Chen, Haoxing Chen, Kun Chen +27
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA…
Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs
Tao Lin, Gaojie Jin, Zongxin Liu +2
Existing vision-language model (VLM) backdoors are usually treated as static vulnerabilities: one-to-one and N-to-N attacks bind one or more triggers to a finite set of targets bef…
Condensing Large-Scale Datasets Directly with Minimal Information Loss
Xinyi Shang, Peng Sun, Bei Shi +2
Recent advancements in scaling dataset distillation rely heavily on decoupled information extraction pipelines, comprising SQUEEZE, RECOVER, and RELABEL stages. Despite their scala…
Self-Adversarial One Step Generation via Condition Shifting
Deyuan Liu, Peng Sun, Yansen Han +3
The push for efficient text to image synthesis has moved the field toward one step sampling, yet existing methods still face a three way tradeoff among fidelity, inference speed, a…
TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
Zhenglin Cheng, Peng Sun, Jianguo Li +1
Recent advances in large multi-modal generative models have demonstrated impressive capabilities in multi-modal generation, including image and video generation. These models are t…
GMem: A Modular Approach for Ultra-Efficient Generative Models
Yi Tang, Peng Sun, Zhenglin Cheng +1
Recent studies indicate that the denoising process in deep generative diffusion models implicitly learns and memorizes semantic information from the data distribution. These findin…