22 citations · 24 across the 4 of their papers we have counts for
4 papers
Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation
Dian Xie, Shitong Shao, Lichen Bai +5
Classifier-free guidance (CFG) has helped diffusion models achieve great conditional generation in various fields. Recently, more diffusion guidance methods have emerged with impro…
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
Dongsheng Wang, Jiequan Cui, Miaoge Li +3
As the open community of large language models (LLMs) matures, multimodal LLMs (MLLMs) have promised an elegant bridge between vision and language. However, current research is inh…
Tuning Multi-mode Token-level Prompt Alignment across Modalities
Dongsheng Wang, Miaoge Li, Xinyang Liu +3
Advancements in prompt tuning of vision-language models have underscored their potential in enhancing open-world visual concept comprehension. However, prior works only primarily f…
Hierarchical Vector Quantized Transformer for Multi-class Unsupervised Anomaly Detection
Ruiying Lu, YuJie Wu, Long Tian +4
Unsupervised image Anomaly Detection (UAD) aims to learn robust and discriminative representations of normal samples. While separate solutions per class endow expensive computation…