Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
Da Zhang, Chenggang Rong, Bingyu Li +4
Large vision-language models (VLMs) have achieved remarkable success in natural scene understanding, yet their application to underwater environments remains largely unexplored. Un…
cs.CV2025
Real-Time Crowd Counting for Embedded Systems with Lightweight Architecture
Zhiyuan Zhao, Yubin Wen, Siyu Yang +3
Crowd counting is a task of estimating the number of the crowd through images, which is extremely valuable in the fields of intelligent security, urban planning, public safety mana…
cs.CV2025
FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation
Bingyu Li, Da Zhang, Zhiyuan Zhao +2
Open-vocabulary segmentation aims to identify and segment specific regions and objects based on text-based descriptions. A common solution is to leverage powerful vision-language m…