most citedUniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

1 citations · 1 across the 7 of their papers we have counts for

collaborators

10 papers

cs.CV2026

From Scale to Speed: Adaptive Test-Time Scaling for Image Editing

Xiangyan Qu, Zhenlong Yuan, Jing Tang +9

Image Chain-of-Thought (Image-CoT) is a test-time scaling paradigm that improves image generation by extending inference time. Most Image-CoT methods focus on text-to-image (T2I) g…

cs.CV2026

Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models

Meiqi Wu, Zhixin Cai, Fufangchen Zhao +13

Video--based world models have emerged along two dominant paradigms: video generation and 3D reconstruction. However, existing evaluation benchmarks either focus narrowly on visual…

cs.CV2026

Video-CoE: Reinforcing Video Event Prediction via Chain of Events

Qile Su, Jing Tang, Rui Chen +2

Despite advances in the application of MLLMs for various video tasks, video event prediction (VEP) remains relatively underexplored. VEP requires the model to perform fine-grained…

cs.RO2026

ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation

Wei Xue, Mingcheng Li, Xuecheng Wu +3

Vision-and-Language Navigation (VLN) requires agents to accurately perceive complex visual environments and reason over navigation instructions and histories. However, existing met…

cs.CV2026

TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward

Yihong Luo, Tianyang Hu, Weijian Luo +1

While few-step generative models have enabled powerful image and video generation at significantly lower cost, generic reinforcement learning (RL) paradigms for few-step models rem…

cs.CL2026

ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning

Xingjian Tao, Yiwei Wang, Yujun Cai +2

Multi-view spatial reasoning remains difficult for current vision-language models. Even when multiple viewpoints are available, models often underutilize cross-view relations and i…