2 papers
cs.RO2025
HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
Yi Li, Yuquan Deng, Jesse Zhang +9
Large foundation models have shown strong open-world generalization to complex problems in vision and language, but similar levels of generalization have yet to be achieved in robo…
cs.CV2025
Magma: A Foundation Model for Multimodal AI Agents
Jianwei Yang, Reuben Tan, Qianhui Wu +10
We present Magma, a foundation model that serves multimodal AI agentic tasks in both the digital and physical worlds. Magma is a significant extension of vision-language (VL) model…