3 papers
cs.CV2025
Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
Dongyang Liu, Peng Gao, David Liu +8
Diffusion model distillation has emerged as a powerful technique for creating efficient few-step and single-step generators. Among these, Distribution Matching Distillation (DMD) a…
cs.MM2025
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
Yaoru Li, Heyu Si, Federico Landi +8
Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing…
cs.CV2025
AesTest: Measuring Aesthetic Intelligence from Perception to Production
Guolong Wang, Heng Huang, Zhiqiang Zhang +3
Perceiving and producing aesthetic judgments is a fundamental yet underexplored capability for multimodal large language models (MLLMs). However, existing benchmarks for image aest…