3 papers
cs.RO2026
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
Chenyv Liu, Wentao Tan, Lei Zhu +4
Standard vision-language-action (VLA) models rely on fitting statistical data priors, limiting their robust understanding of underlying physical dynamics. Reinforcement learning en…
cs.CV2025
Unified modality separation: A vision-language framework for unsupervised domain adaptation
Xinyao Li, Jingjing Li, Zhekai Du +2
Unsupervised domain adaptation (UDA) enables models trained on a labeled source domain to handle new unlabeled domains. Recently, pre-trained vision-language models (VLMs) have dem…
cs.CV2025
Generalizing vision-language models to novel domains: A comprehensive survey
Xinyao Li, Jingjing Li, Fengling Li +3
Recently, vision-language pretraining has emerged as a transformative technique that integrates the strengths of both visual and textual modalities, resulting in powerful vision-la…