1 citations · 1 across the 2 of their papers we have counts for
4 papers · 1 filter
From Editor to Dense Geometry Estimator
JiYuan Wang, Chunyu Lin, Lei Sun +5
Leveraging visual priors from pre-trained text-to-image (T2I) generative models has shown success in dense prediction. However, dense prediction is inherently an image-to-image tas…
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
Sule Bai, Mingxing Li, Yong Liu +5
Traditional visual grounding methods primarily focus on single-image scenarios with simple textual references. However, extending these methods to real-world scenarios that involve…
FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
Rui Lan, Yancheng Bai, Xu Duan +6
Scene text editing aims to modify or add texts on images while ensuring text fidelity and overall visual quality consistent with the background. Recent methods are primarily built…
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
Mingxing Li, Rui Wang, Lei Sun +2
The rapid expansion of mobile internet has resulted in a substantial increase in user-generated content (UGC) images, thereby making the thorough assessment of UGC images both urge…