3 papers
cs.CV2025
Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding
Yuzhen Li, Min Liu, Zhaoyang Li +4
Monocular 3D Visual Grounding (Mono3DVG) is an emerging task that locates 3D objects in RGB images using text descriptions with geometric cues. However, existing methods face two k…
cs.CV2025
A Dual-stage Prompt-driven Privacy-preserving Paradigm for Person Re-Identification
Ruolin Li, Min Liu, Yuan Bian +4
With growing concerns over data privacy, researchers have started using virtual data as an alternative to sensitive real-world images for training person re-identification (Re-ID)…
cs.CV2025
Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding
Yuzhen Li, Min Liu, Yuan Bian +4
Monocular 3D visual grounding is a novel task that aims to locate 3D objects in RGB images using text descriptions with explicit geometry information. Despite the inclusion of geom…