2 papers
cs.RO2026
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents
Hao Sun, Yu Song, Shiyu Teng +2
VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current single-frame architectures suff…
cs.CV2025
EPIC: Efficient Prompt Interaction for Text-Image Classification
Xinyao Yu, Hao Sun, Zeyu Ling +5
In recent years, large-scale pre-trained multimodal models (LMMs) generally emerge to integrate the vision and language modalities, achieving considerable success in multimodal tas…