activity
20242026
collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Understanding Degradation with Vision Language Model

Guanzhou Lan, Chenyi Liao, Yuqi Yang +5

Understanding visual degradations is a critical yet challenging problem in computer vision. While recent Vision-Language Models (VLMs) excel at qualitative description, they often…

cs.CV2025

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

Pingrui Zhang, Yifei Su, Pengyuan Wu +7

Vision-and-Language Navigation (VLN) requires the agent to navigate by following natural instructions under partial observability, making it difficult to align perception with lang…

cs.CV2025

AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations

Junli Liu, Qizhi Chen, Zhigang Wang +6

Visual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual groundin…

cs.CV2025

Exploring the Potential of Encoder-free Architectures in 3D LMMs

Yiwen Tang, Zoey Guo, Zhuhao Wang +8

Encoder-free architectures have been preliminarily explored in the 2D Large Multimodal Models (LMMs), yet it remains an open question whether they can be effectively applied to 3D…

cs.CV2024

Night-to-Day Translation via Illumination Degradation Disentanglement

Guanzhou Lan, Yuqi Yang, Zhigang Wang +3

Night-to-Day translation (Night2Day) aims to achieve day-like vision for nighttime scenes. However, processing night images with complex degradations remains a significant challeng…

cs.CV2024

Open-Vocabulary Octree-Graph for 3D Scene Understanding

Zhigang Wang, Yifei Su, Chenhui Li +4

Open-vocabulary 3D scene understanding is indispensable for embodied agents. Recent works leverage pretrained vision-language models (VLMs) for object segmentation and project them…