6 papers
ALADIN:Attribute-Language Distillation Network for Person Re-Identification
Wang Zhou, Boran Duan, Haojun Ai +2
Recent vision-language models such as CLIP provide strong cross-modal alignment, but current CLIP-guided ReID pipelines rely on global features and fixed prompts. This limits their…
Semantic Context Matters: Improving Conditioning for Autoregressive Models
Dongyang Jin, Ryan Xu, Jianhao Zeng +4
Recently, autoregressive (AR) models have shown strong potential in image generation, offering better scalability and easier integration with unified multi-modal systems compared t…
TEXTS-Diff: TEXTS-Aware Diffusion Model for Real-World Text Image Super-Resolution
Haodong He, Xin Zhan, Yancheng Bai +3
Real-world text image super-resolution aims to restore overall visual quality and text legibility in images suffering from diverse degradations and text distortions. However, the s…
FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
Rui Lan, Yancheng Bai, Xu Duan +6
Scene text editing aims to modify or add texts on images while ensuring text fidelity and overall visual quality consistent with the background. Recent methods are primarily built…
SCALAR: Scale-wise Controllable Visual Autoregressive Learning
Ryan Xu, Dongyang Jin, Yancheng Bai +4
Controllable image synthesis, which enables fine-grained control over generated outputs, has emerged as a key focus in visual generative modeling. However, controllable generation…
RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
Haodong He, Yancheng Bai, Rui Lan +4
The rich textual information of large vision-language models (VLMs) combined with the powerful generative prior of pre-trained text-to-image (T2I) diffusion models has achieved imp…