3 papers
cs.CV2026
DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering
Yue Zhang, Xiangyu Li, Wanshu Fan +2
Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered si…
cs.CV2026
MambaLIE: Scene Light Intensity-Boosted Low-Light Image Enhancement with State Space Model
Wanshu Fan, Xiangyu Li, Cong Wang +4
Images captured by consumer electronic devices, such as mobile phones and digital cameras, often suffer from low-light degradation due to sensor limitations and imaging pipelines,…
cs.CV2025
Iterative Optimal Attention and Local Model for Single Image Rain Streak Removal
Xiangyu Li, Wanshu Fan, Yue Shen +5
High-fidelity imaging is crucial for the successful safety supervision and intelligent deployment of vision-based measurement systems (VBMS). It ensures high-quality imaging in VBM…