works on

From the 1 of 18 linked papers with an AI index.

activity
20242026
most citedQFT: Quantized Full-parameter Tuning of LLMs with Affordable Resources

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

Post-Training Pruning for Diffusion Transformers

Chengzhi Hu, Xuewen Liu, Jing Zhang +3

The paper introduces DiT-Pruning, a post‑training pruning method tailored for Diffusion Transformers that uses a new energy‑based saliency metric and clustering‑aware granularity t…

cs.CV2026

Arena as Offline Reward: Efficient Fine-Grained Preference Optimization for Diffusion Models

Zhikai Li, Yue Zhao, Edward Zhongwei Zhang +4

Reinforcement learning from human feedback (RLHF) effectively promotes preference alignment of text-to-image (T2I) diffusion models. To improve computational efficiency, direct pre…

cs.CV2026

K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge

Zhikai Li, Jiatong Li, Xuewen Liu +7

The rapid development of visual generative models raises the need for more scalable and human-aligned evaluation methods. While the crowdsourced Arena platforms offer human prefere…

cs.CV2026

PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models

Xuewen Liu, Zhikai Li, Jing Zhang +2

AutoRegressive Visual Generation (ARVG) models retain an architecture compatible with language models, while achieving performance comparable to diffusion-based models. Quantizatio…

cs.CV2025

Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation

Xuewen Liu, Zhikai Li, Jing Zhang +2

Diffusion Transformers dominate video generation, but the quadratic complexity of attention computation introduces substantial latency. Attention sparsity reduces computational cos…

cs.CV2025

LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation

Lianwei Yang, Haokun Lin, Tianchen Zhao +6

Diffusion Transformers (DiTs) have achieved impressive performance in text-to-image and text-to-video generation. However, their high computational cost and large parameter sizes p…