6 papers
RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models
Qihui Zhu, Yuchen Wang, Zijian Wen +7
On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories generated by the student. Howe…
Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning
Tao Zhang, Qixuan Fan, Yiyuan Liang +7
Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is effective, it raises concern…
Factor-Aware Mixture-of-Experts with Pretrained Encoder for Combinatorial Generalization
Feihong Zhang, Guojian Zhan, Zeyu He +8
The integration of pretrained encoders with diffusion policies has become a dominant paradigm for visual robotic manipulation. However, it still struggles to generalize across comp…
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models
Yueyi Sun, Yuhao Wang, Jason Li +8
Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limi…
Determinism of Randomness: Prompt-Residual Seed Shaping for Diffusion Generation
Song Yan, Wei Zhai, Chenfeng Wang +9
Diffusion models start generation from an isotropic Gaussian latent, yet changing only the random seed can lead to large differences in prompt faithfulness, composition, and visual…
Break Stylistic Sophon: Are We Really Meant to Confine the Imagination in Style Transfer?
Gary Song Yan, Yusen Zhang, Jinyu Zhao +9
In this pioneering study, we introduce StyleWallfacer, a groundbreaking unified training and inference framework, which not only addresses various issues encountered in the style t…