4 papers
Grounded Reinforcement Learning for Visual Reasoning
Gabriel Sarch, Snigdha Saha, Naitik Khandelwal +4
While reinforcement learning (RL) over chains of thought has significantly advanced language models in tasks such as mathematics and coding, visual reasoning introduces added compl…
Elastic Attention Cores for Scalable Vision Transformers
Alan Z. Song, Yinjie Chen, Mu Nan +8
Vision Transformers (ViTs) achieve strong data-driven scaling by leveraging all-to-all self-attention. However, this flexibility incurs a computational cost that scales quadratical…
VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
Gabriel Sarch, Lawrence Jang, Michael J. Tarr +3
Large-scale generative language and vision-language models (LLMs and VLMs) excel in few-shot learning but require high-quality demonstrations. We propose In-Context Abstraction Lea…
Reanimating Images using Neural Representations of Dynamic Stimuli
Jacob Yeung, Andrew F. Luo, Gabriel Sarch +3
While computer vision models have made incredible strides in static image recognition, they still do not match human performance in tasks that require the understanding of complex,…