3 papers
cs.NI2026
MapViT: A Two-Stage ViT-Based Framework for Real-Time Radio Quality Map Prediction in Dynamic Environments
Cyril Shih-Huan Hsu, Xi Li, Lanfranco Zanzi +3
Recent advancements in mobile and wireless networks are unlocking the full potential of robotic autonomy, enabling robots to take advantage of ultra-low latency, high data throughp…
cs.CV2025
Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
Chaoxiang Cai, Longrong Yang, Minghe Weng +3
The mixture-of-experts (MoE) architecture, which replaces dense networks with sparse ones, has attracted significant attention in large vision-language models (LVLMs) for achieving…
cs.MM2025
Mitigating Image Captioning Hallucinations in Vision-Language Models
Fei Zhao, Chengcui Zhang, Runlin Zhang +2
Hallucinations in vision-language models (VLMs) hinder reliability and real-world applicability, usually stemming from distribution shifts between pretraining data and test samples…