works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.CV2026

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding

Kai Chen, Ming Dai, Wenxuan Cheng +1

The paper introduces ScanFocus, a coarse-to-fine framework for spatio-temporal video grounding that first scans videos globally to generate coarse object proposals and then refines…

cs.CV2025

DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy

Ming Dai, Wenxuan Cheng, Jiang-jiang Liu +4

Referring Image Segmentation (RIS) is a challenging task that aims to segment objects in an image based on natural language expressions. While prior studies have predominantly conc…

physics.flu-dyn2025

Effective transport by 2D turbulence: Vortex-gas theory vs. scale-invariant inverse cascade

Julie Meunier, Basile Gallet

The scale-invariant inverse energy cascade is a hallmark of 2D turbulence, with its theoretical energy spectrum observed in both direct numerical simulations (DNS) and laboratory e…

cs.CV2025

SpatialBot: Precise Spatial Understanding with Vision Language Models

Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan +4

Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding, however they are still struggling with spatial understanding which is the foundation o…

cs.CV2025

Precise GPS-Denied UAV Self-Positioning via Context-Enhanced Cross-View Geo-Localization

Yuanze Xu, Ming Dai, Wenxiao Cai +1

Image retrieval has been employed as a robust complementary technique to address the challenge of Unmanned Aerial Vehicles (UAVs) self-positioning. However, most existing methods p…

cs.CV2024

Object-level Geometric Structure Preserving for Natural Image Stitching

Wenxiao Cai, Wankou Yang

The topic of stitching images with globally natural structures holds paramount significance, with two main goals: pixel-level alignment and distortion prevention. The existing appr…