most citedUniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

1 citations · 1 across the 7 of their papers we have counts for

collaborators

10 papers

cs.CV2026

TEXTS-Diff: TEXTS-Aware Diffusion Model for Real-World Text Image Super-Resolution

Haodong He, Xin Zhan, Yancheng Bai +3

Real-world text image super-resolution aims to restore overall visual quality and text legibility in images suffering from diverse degradations and text distortions. However, the s…

cs.RO2025

Real-world Reinforcement Learning from Suboptimal Interventions

Yinuo Zhao, Huiqian Jin, Lechun Jiang +9

Real-world reinforcement learning (RL) offers a promising approach to training precise and dexterous robotic manipulation policies in an online manner, enabling robots to learn fro…

cs.CV2025

RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution

Haodong He, Yancheng Bai, Rui Lan +4

The rich textual information of large vision-language models (VLMs) combined with the powerful generative prior of pre-trained text-to-image (T2I) diffusion models has achieved imp…

cs.CV2025

RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space

Jingyun Liang, Jingkai Zhou, Shikai Li +5

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control ov…

cs.CV2025

SCALAR: Scale-wise Controllable Visual Autoregressive Learning

Ryan Xu, Dongyang Jin, Yancheng Bai +4

Controllable image synthesis, which enables fine-grained control over generated outputs, has emerged as a key focus in visual generative modeling. However, controllable generation…

cs.CV20251 cited

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

Sule Bai, Mingxing Li, Yong Liu +5

Traditional visual grounding methods primarily focus on single-image scenarios with simple textual references. However, extending these methods to real-world scenarios that involve…