3 papers
cs.RO2026
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
Wenda Yu, Tianshi Wang, Fengling Li +2
Vision-Language-Action (VLA) models have demonstrated strong performance in robotic manipulation, yet their closed-loop deployment is hindered by the high latency and compute cost…
cs.CV2025
Generalizing vision-language models to novel domains: A comprehensive survey
Xinyao Li, Jingjing Li, Fengling Li +3
Recently, vision-language pretraining has emerged as a transformative technique that integrates the strengths of both visual and textual modalities, resulting in powerful vision-la…
cs.IR2024
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
Tianshi Wang, Fengling Li, Lei Zhu +3
With the exponential surge in diverse multi-modal data, traditional uni-modal retrieval methods struggle to meet the needs of users seeking access to data across various modalities…