5 papers
DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering
Yue Zhang, Xiangyu Li, Wanshu Fan +2
Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered si…
MambaLIE: Scene Light Intensity-Boosted Low-Light Image Enhancement with State Space Model
Wanshu Fan, Xiangyu Li, Cong Wang +4
Images captured by consumer electronic devices, such as mobile phones and digital cameras, often suffer from low-light degradation due to sensor limitations and imaging pipelines,…
View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
Yuanyuan Liu, Haiyang Mei, Dongyang Zhan +4
3D visual grounding (3DVG) identifies objects in 3D scenes from language descriptions. Existing zero-shot approaches leverage 2D vision-language models (VLMs) by converting 3D spat…
Iterative Optimal Attention and Local Model for Single Image Rain Streak Removal
Xiangyu Li, Wanshu Fan, Yue Shen +5
High-fidelity imaging is crucial for the successful safety supervision and intelligent deployment of vision-based measurement systems (VBMS). It ensures high-quality imaging in VBM…
Semantic-Guided Global-Local Collaborative Networks for Lightweight Image Super-Resolution
Wanshu Fan, Yue Wang, Cong Wang +3
Single-Image Super-Resolution (SISR) plays a pivotal role in enhancing the accuracy and reliability of measurement systems, which are integral to various vision-based instrumentati…