7 citations · 8 across the 3 of their papers we have counts for
3 papers
cs.CV2025★ 1 cited
MC-VTON: Minimal Control Virtual Try-On Diffusion Transformer
Junsheng Luan, Guangyuan Li, Lei Zhao +1
Virtual try-on methods based on diffusion models achieve realistic try-on effects. They use an extra reference network or an additional image encoder to process multiple conditiona…
cs.CV2024★ 7 cited
CogVLM2: Visual Language Models for Image and Video Understanding
Wenyi Hong, Weihan Wang, Ming Ding +22
Beginning with VisualGLM and CogVLM, we are continuously exploring VLMs in pursuit of enhanced vision-language fusion, efficient higher-resolution architecture, and broader modalit…
cs.LG2024
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
Xingwu Chen, Lei Zhao, Difan Zou
Despite the remarkable success of transformer-based models in various real-world tasks, their underlying mechanisms remain poorly understood. Recent studies have suggested that tra…