14 papers
Latent Implicit Visual Reasoning
Kelvin Li, Chuyi Shang, Leonid Karlinsky +3
While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning modality. As a result, they are l…
Learning to Grasp Anything by Playing with Random Toys
Dantong Niu, Yuvan Sharma, Baifeng Shi +11
Robotic manipulation policies often struggle to generalize to novel objects, limiting their real-world utility. In contrast, cognitive science suggests that children develop genera…
EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
Ruijie Zheng, Dantong Niu, Yuqi Xie +12
Human behavior is among the most scalable sources of data for learning physical intelligence, yet how to effectively leverage it for dexterous manipulation remains unclear. While p…
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
Brandon Huang, Hang Hua, Zhuoran Yu +3
While Vision-language models (VLMs) have demonstrated remarkable performance across multi-modal tasks, their choice of vision encoders presents a fundamental weakness: their low-le…
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
Chancharik Mitra, Yusen Luo, Raj Saravanan +7
Vision-Language Action (VLAs) models promise to extend the remarkable success of vision-language models (VLMs) to robotics. Yet, unlike VLMs in the vision-language domain, VLAs for…
Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs
Nasim Borazjanizadeh, Roei Herzig, Eduard Oks +3
Human reasoning relies on constructing and manipulating mental models -- simplified internal representations of situations used to understand and solve problems. Conceptual diagram…