From the 1 of 18 linked papers with an AI index.
1 citations · 1 across the 6 of their papers we have counts for
12 papers · 1 filter
Post-Training Pruning for Diffusion Transformers
Chengzhi Hu, Xuewen Liu, Jing Zhang +3
The paper introduces DiT-Pruning, a post‑training pruning method tailored for Diffusion Transformers that uses a new energy‑based saliency metric and clustering‑aware granularity t…
Arena as Offline Reward: Efficient Fine-Grained Preference Optimization for Diffusion Models
Zhikai Li, Yue Zhao, Edward Zhongwei Zhang +4
Reinforcement learning from human feedback (RLHF) effectively promotes preference alignment of text-to-image (T2I) diffusion models. To improve computational efficiency, direct pre…
K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge
Zhikai Li, Jiatong Li, Xuewen Liu +7
The rapid development of visual generative models raises the need for more scalable and human-aligned evaluation methods. While the crowdsourced Arena platforms offer human prefere…
PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models
Xuewen Liu, Zhikai Li, Jing Zhang +2
AutoRegressive Visual Generation (ARVG) models retain an architecture compatible with language models, while achieving performance comparable to diffusion-based models. Quantizatio…
Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation
Xuewen Liu, Zhikai Li, Jing Zhang +2
Diffusion Transformers dominate video generation, but the quadratic complexity of attention computation introduces substantial latency. Attention sparsity reduces computational cos…
LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
Lianwei Yang, Haokun Lin, Tianchen Zhao +6
Diffusion Transformers (DiTs) have achieved impressive performance in text-to-image and text-to-video generation. However, their high computational cost and large parameter sizes p…