5 papers
Colosseum V2: Benchmarking Generalization for Vision Language Action Models
Jeremy Morgan, Prajwal Vijay, Hyeonho Oh +6
Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language pre-training. This progress…
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
Ishika Singh, Ankit Goyal, Stan Birchfield +3
We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware…
PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language Models
Wang Bill Zhu, Miaosen Chai, Ishika Singh +2
We propose PSALM-V, the first autonomous neuro-symbolic learning system able to induce symbolic action semantics (i.e., pre- and post-conditions) in visual environments through int…
TwoStep: Multi-agent Task Planning using Classical Planners and Large Language Models
David Bai, Ishika Singh, David Traum +1
Classical planning formulations like the Planning Domain Definition Language (PDDL) admit action sequences guaranteed to achieve a goal state given an initial state if any are poss…
Language Models can Infer Action Semantics for Symbolic Planners from Environment Feedback
Wang Zhu, Ishika Singh, Robin Jia +1
Symbolic planners can discover a sequence of actions from initial to goal states given expert-defined, domain-specific logical action semantics. Large Language Models (LLMs) can di…