111 citations · 216 across the 29 of their papers we have counts for
12 papers · 1 filter
Visual Representations inside the Language Model
Benlin Liu, Amita Kamath, Madeleine Grunde-McLaughlin +2
Despite interpretability work analyzing VIT encoders and transformer activations, we don't yet understand why Multimodal Language Models (MLMs) struggle on perception-heavy tasks.…
MolmoAct: Action Reasoning Models that can Reason in Space
Jason Lee, Jiafei Duan, Haoquan Fang +16
Reasoning is central to purposeful action, yet most robotic foundation models map perception and instructions directly to control, which limits adaptability, generalization, and se…
Rethinking Human Preference Evaluation of LLM Rationales
Ziang Li, Manasi Ganti, Zixian Ma +3
Large language models (LLMs) often generate natural language rationales -- free-form explanations that help improve performance on complex reasoning tasks and enhance interpretabil…
ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
Ge Yan, Jiyue Zhu, Yuquan Deng +8
This paper introduces ManiFlow, a visuomotor imitation learning policy for general robot manipulation that generates precise, high-dimensional actions conditioned on diverse visual…
Reinforced Visual Perception with Tools
Zetong Zhou, Dongping Chen, Zixian Ma +6
Visual reasoning, a cornerstone of human intelligence, encompasses complex perceptual and logical processes essential for solving diverse visual problems. While advances in compute…
MultiRef: Controllable Image Generation with Multiple Visual References
Ruoxi Chen, Dongping Chen, Siyuan Wu +6
Visual designers naturally draw inspiration from multiple visual references, combining diverse elements and aesthetic principles to create artwork. However, current image generativ…