5 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.RO2024
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
Dantong Niu, Yuvan Sharma, Giscard Biamby +5
In recent years, instruction-tuned Large Multimodal Models (LMMs) have been successful at several tasks, including image captioning and visual question answering; yet leveraging th…
cs.RO2024★ 5 cited
Humanoid Locomotion as Next Token Prediction
Ilija Radosavovic, Bike Zhang, Baifeng Shi +5
We cast real-world humanoid control as a next token prediction problem, akin to predicting the next word in language. Our model is a causal transformer trained via autoregressive p…
cs.CV2023★ 2 cited
Top-Down Visual Attention from Analysis by Synthesis
Baifeng Shi, Trevor Darrell, Xin Wang
Current attention algorithms (e.g., self-attention) are stimulus-driven and highlight all the salient objects in an image. However, intelligent agents like humans often guide their…