From the 1 of 6 linked papers with an AI index.
1 citations · 1 across the 3 of their papers we have counts for
6 papers
VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
Yupeng Zheng, Kai Zou, Bin Liu +1
VisCo introduces a training-efficient self-compression framework that reuses a pretrained vision-language model as an intrinsic autoencoder to compress visual tokens into a small s…
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
Dian Zheng, Manyuan Zhang, Hongyu Li +4
Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi-task training. Due to inherent tas…
Advancing Aesthetic Image Generation via Composition Transfer
Kai Zou, Zhiwei Zhao, Bin Liu +1
Composition is a cornerstone of visual aesthetics, influencing the appeal of an image. While its principles operate independently of specific content, in practice, composition is o…
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
Kai Zou, Dian Zheng, Hongbo Liu +3
Autoregressive (AR) diffusion offers a promising framework for generating videos of theoretically infinite length. However, a major challenge is maintaining temporal continuity whi…
Normalized Attention Guidance: Universal Negative Guidance for Diffusion Models
Dar-Yen Chen, Hmrishav Bandyopadhyay, Kai Zou +1
Negative guidance -- explicitly suppressing unwanted attributes -- remains a fundamental challenge in diffusion models, particularly in few-step sampling regimes. While Classifier-…
NitroFusion: High-Fidelity Single-Step Diffusion through Dynamic Adversarial Training
Dar-Yen Chen, Hmrishav Bandyopadhyay, Kai Zou +1
We introduce NitroFusion, a fundamentally different approach to single-step diffusion that achieves high-quality generation through a dynamic adversarial framework. While one-step…