2 papers
cs.CV2026
Skill-Conditioned Visual Geolocation for Vision-Language Models
Chenjie Yang, Yutian Jiang, Yutong Deng +1
Vision-language models (VLMs) have shown a promising ability in image geolocation, but they still lack structured geographic reasoning and the capacity for autonomous self-evolutio…
cs.LG2026
Can Explicit Physical Feasibility Benefit VLA Learning? An Empirical Study
Yubai Wei, Chen Wu, Hashem Haghbayan
Vision-Language-Action (VLA) models map multimodal inputs directly to robot actions and are typically trained through large-scale imitation learning. While this paradigm has shown…