3 papers
cs.CV2026
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers
Guozhen Zhang, Xuerui Qiu, Yutao Cui +11
Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDR…
cs.CV2025
Differentiable Solver Search for Fast Diffusion Sampling
Shuai Wang, Zexian Li, Qipeng zhang +5
Diffusion models have demonstrated remarkable generation quality but at the cost of numerous function evaluations. Recently, advanced ODE-based solvers have been developed to mitig…
cs.CV2025
DMM: Building a Versatile Image Generation Model via Distillation-Based Model Merging
Tianhui Song, Weixin Feng, Shuai Wang +4
The success of text-to-image (T2I) generation models has spurred a proliferation of numerous model checkpoints fine-tuned from the same base model on various specialized datasets.…