activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

Jiahao Wen, Hang Yu, Zhedong Zheng

Text-guided drone geo-localization aims to identify a target region in a large-scale image gallery from a natural-language description. Existing methods mainly formulate this task…

cs.CV2025

WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization

Jiahao Wen, Hang Yu, Zhedong Zheng

Visual geo-localization for drones faces critical degradation under weather perturbations, \eg, rain and fog, where existing methods struggle with two inherent limitations: 1) Heav…

cs.CV2025

CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval

Hang Yu, Jiahao Wen, Zhedong Zheng

Text-based person retrieval aims to identify specific individuals within an image database using textual descriptions. Due to the high cost of annotation and privacy protection, re…

cs.CV2024

Harnessing Uncertainty-aware Bounding Boxes for Unsupervised 3D Object Detection

Ruiyang Zhang, Hu Zhang, Hang Yu +1

Unsupervised 3D object detection aims to identify objects of interest from unlabeled raw data, such as LiDAR points. Recent approaches usually adopt pseudo 3D bounding boxes (3D bb…

cs.CV2024

Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene

Ruiyang Zhang, Hu Zhang, Hang Yu +1

The unsupervised 3D object detection is to accurately detect objects in unstructured environments with no explicit supervisory signals. This task, given sparse LiDAR point clouds,…