3 papers
cs.RO2025
Tri-Select: A Multi-Stage Visual Data Selection Framework for Mobile Visual Crowdsensing
Jiayu Zhang, Kaixing Zhao, Tianhao Shao +2
Mobile visual crowdsensing enables large-scale, fine-grained environmental monitoring through the collection of images from distributed mobile devices. However, the resulting data…
cs.RO2025
LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
Zeyi Kang, Liang He, Yanxin Zhang +2
Multimodal semantic learning plays a critical role in embodied intelligence, especially when robots perceive their surroundings, understand human instructions, and make intelligent…
cs.RO2025
M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer
Yanxin Zhang, Liang He, Zeyi Kang +2
In recent years, multimodal learning has become essential in robotic vision and information fusion, especially for understanding human behavior in complex environments. However, cu…