3 citations · 6 across the 2 of their papers we have counts for
3 papers
cs.CV2024★ 3 cited
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
Yang Zhan, Zhitong Xiong, Yuan Yuan
Large language models (LLMs) have recently been extended to the vision-language realm, obtaining impressive general multi-modal capabilities. However, the exploration of multi-moda…
cs.CV2023
Mono3DVG: 3D Visual Grounding in Monocular Images
Yang Zhan, Yuan Yuan, Zhitong Xiong
We introduce a novel task of 3D visual grounding in monocular RGB images using language descriptions with both appearance and geometry information. Specifically, we build a large-s…
cs.CV2021★ 3 cited
Deep Superpixel-based Network for Blind Image Quality Assessment
Guangyi Yang, Yang Zhan., Yuxuan Wang
The goal in a blind image quality assessment (BIQA) model is to simulate the process of evaluating images by human eyes and accurately assess the quality of the image. Although man…