1 citations · 1 across the 7 of their papers we have counts for
6 papers · 1 filter
Accelerating Controllable Generation via Hybrid-grained Cache
Lin Liu, Huixia Ben, Shuo Wang +4
Controllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation…
Accelerating Diffusion Transformer via Gradient-Optimized Cache
Junxiang Qiu, Lin Liu, Shuo Wang +3
Feature caching has emerged as an effective strategy to accelerate diffusion transformer (DiT) sampling through temporal feature reuse. It is a challenging problem since (1) Progre…
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
Jinda Lu, Junkang Wu, Jinghan Li +6
Direct Preference Optimization (DPO) has shown effectiveness in aligning multi-modal large language models (MLLM) with human preferences. However, existing methods exhibit an imbal…
Accelerating Diffusion Transformer via Error-Optimized Cache
Junxiang Qiu, Shuo Wang, Jinda Lu +4
Diffusion Transformer (DiT) is a crucial method for content generation. However, it needs a lot of time to sample. Many studies have attempted to use caching to reduce the time con…
Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation
Jinda Lu, Shuo Wang, Yanbin Hao +3
Recent adaptations can boost the low-shot capability of Contrastive Vision-Language Pre-training (CLIP) by effectively facilitating knowledge transfer. However, these adaptation me…
Boosting Few-Shot Learning via Attentive Feature Regularization
Xingyu Zhu, Shuo Wang, Jinda Lu +3
Few-shot learning (FSL) based on manifold regularization aims to improve the recognition capacity of novel objects with limited training samples by mixing two samples from differen…