3 citations · 3 across the 2 of their papers we have counts for
3 papers
cs.CV2026
UAVBench and UAVIT-1M: Benchmarking and Enhancing MLLMs for Low-Altitude UAV Vision-Language Understanding
Yang Zhan, Yuan Yuan
Multimodal Large Language Models (MLLMs) have made significant strides in natural images and satellite remote sensing images. However, understanding low-altitude drone scenarios re…
cs.CV2024★ 3 cited
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
Yang Zhan, Zhitong Xiong, Yuan Yuan
Large language models (LLMs) have recently been extended to the vision-language realm, obtaining impressive general multi-modal capabilities. However, the exploration of multi-moda…
cs.CV2023
Mono3DVG: 3D Visual Grounding in Monocular Images
Yang Zhan, Yuan Yuan, Zhitong Xiong
We introduce a novel task of 3D visual grounding in monocular RGB images using language descriptions with both appearance and geometry information. Specifically, we build a large-s…