1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
AddressCLIP: Empowering Vision-Language Models for City-wide Image Address Localization
Shixiong Xu, Chenghao Zhang, Lubin Fan +3
In this study, we introduce a new problem raised by social media and photojournalism, named Image Address Localization (IAL), which aims to predict the readable textual address whe…
cs.CV2024
Reusable Architecture Growth for Continual Stereo Matching
Chenghao Zhang, Gaofeng Meng, Bin Fan +4
The remarkable performance of recent stereo depth estimation models benefits from the successful use of convolutional neural networks to regress dense disparity. Akin to most tasks…
cs.CV2024★ 1 cited
Enhancing Visual Continual Learning with Language-Guided Supervision
Bolin Ni, Hongbo Zhao, Chenghao Zhang +4
Continual learning (CL) aims to empower models to learn new tasks without forgetting previously acquired knowledge. Most prior works concentrate on the techniques of architectures,…