activity
20242026
collaborators

14 papers

cs.CV2026

Latent Implicit Visual Reasoning

Kelvin Li, Chuyi Shang, Leonid Karlinsky +3

While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning modality. As a result, they are l…

cs.RO2026

Learning to Grasp Anything by Playing with Random Toys

Dantong Niu, Yuvan Sharma, Baifeng Shi +11

Robotic manipulation policies often struggle to generalize to novel objects, limiting their real-world utility. In contrast, cognitive science suggests that children develop genera…

cs.RO2026

EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data

Ruijie Zheng, Dantong Niu, Yuqi Xie +12

Human behavior is among the most scalable sources of data for learning physical intelligence, yet how to effectively leverage it for dexterous manipulation remains unclear. While p…

cs.CV2025

DAVE: A VLM Vision Encoder for Document Understanding and Web Agents

Brandon Huang, Hang Hua, Zhuoran Yu +3

While Vision-language models (VLMs) have demonstrated remarkable performance across multi-modal tasks, their choice of vision encoders presents a fundamental weakness: their low-le…

cs.RO2025

Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations

Chancharik Mitra, Yusen Luo, Raj Saravanan +7

Vision-Language Action (VLAs) models promise to extend the remarkable success of vision-language models (VLMs) to robotics. Yet, unlike VLMs in the vision-language domain, VLAs for…

cs.AI2025

Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs

Nasim Borazjanizadeh, Roei Herzig, Eduard Oks +3

Human reasoning relies on constructing and manipulating mental models -- simplified internal representations of situations used to understand and solve problems. Conceptual diagram…