activity
20242026
collaborators

5 papers

cs.CV2026

DSFC-Net: A Dual-Encoder Spatial and Frequency Co-Awareness Network for Rural Road Extraction

Zhengbo Zhang, Yihe Tian, Wanke Xia +6

Accurate extraction of rural roads from high-resolution remote sensing imagery is essential for infrastructure planning and sustainable development. However, this task presents uni…

cs.CV2025

Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering

Yuyang Hong, Jiaqi Gu, Qi Yang +6

Knowledge-based visual question answering (KB-VQA) requires visual language models (VLMs) to integrate visual understanding with external knowledge retrieval. Although retrieval-au…

cs.CV2025

Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger

Qi Yang, Chenghao Zhang, Lubin Fan +3

Recent advancements in Large Vision Language Models (LVLMs) have significantly improved performance in Visual Question Answering (VQA) tasks through multimodal Retrieval-Augmented…

cs.CV2025

Efficient Redundancy Reduction for Open-Vocabulary Semantic Segmentation

Lin Chen, Qi Yang, Kun Ding +5

Open-vocabulary semantic segmentation (OVSS) is an open-world task that aims to assign each pixel within an image to a specific class defined by arbitrary text descriptions. While…

cs.CL2024

Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature

Lingdong Shen, Qigqi, Kun Ding +2

Scientific Literature charts often contain complex visual elements, including multi-plot figures, flowcharts, structural diagrams and etc. Evaluating multimodal models using these…