3 papers
cs.CV2025
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
Yiming Chen, Junlin Han, Tianyi Bai +3
While Multimodal Large Language Models (MLLMs) are adept at answering what is in an image-identifying objects and describing scenes-they often lack the ability to understand how an…
cs.LG2025
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
Junlin Han, Shengbang Tong, David Fan +4
Large Language Models (LLMs), despite being trained on text alone, surprisingly develop rich visual priors. These priors allow latent visual capabilities to be unlocked for vision…
cs.RO2025
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Irving Fang, Juexiao Zhang, Shengbang Tong +1
One promise that Vision-Language-Action (VLA) models hold over traditional imitation learning for robotics is to leverage the broad generalization capabilities of large Vision-Lang…