Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
Quoc-Huy Trinh, Mustapha Abdullahi, Bo Zhao +1
Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in vision-language understanding, yet their high computational cost limits deployment i…
cs.CV2025
Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
Quoc-Huy Trinh
Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in vision-language understanding, yet their high computational cost limits deployment i…
cs.CV2025
Evaluating Deep Learning Models for African Wildlife Image Classification: From DenseNet to Vision Transformers
Lukman Jibril Aliyu, Umar Sani Muhammad, Bilqisu Ismail +5
Wildlife populations in Africa face severe threats, with vertebrate numbers declining by over 65% in the past five decades. In response, image classification using deep learning ha…