Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Group Relative Attention Guidance for Image Editing
Xuanpu Zhang, Xuesong Niu, Ruidong Chen +6
Recently, image editing based on Diffusion-in-Transformer models has undergone rapid development. However, existing editing methods often lack effective control over the degree of…
cs.CV2024
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
Yang Jin, Kun Xu, Liwei Chen +12
Recently, the remarkable advance of the Large Language Model (LLM) has inspired researchers to transfer its extraordinary reasoning capability to both vision and language data. How…