activity
20242026
collaborators

5 papers

cs.RO2026

Colosseum V2: Benchmarking Generalization for Vision Language Action Models

Jeremy Morgan, Prajwal Vijay, Hyeonho Oh +6

Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language pre-training. This progress…

cs.RO2025

OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model

Ishika Singh, Ankit Goyal, Stan Birchfield +3

We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware…

cs.RO2025

PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language Models

Wang Bill Zhu, Miaosen Chai, Ishika Singh +2

We propose PSALM-V, the first autonomous neuro-symbolic learning system able to induce symbolic action semantics (i.e., pre- and post-conditions) in visual environments through int…

cs.AI2025

TwoStep: Multi-agent Task Planning using Classical Planners and Large Language Models

David Bai, Ishika Singh, David Traum +1

Classical planning formulations like the Planning Domain Definition Language (PDDL) admit action sequences guaranteed to achieve a goal state given an initial state if any are poss…

cs.AI2024

Language Models can Infer Action Semantics for Symbolic Planners from Environment Feedback

Wang Zhu, Ishika Singh, Robin Jia +1

Symbolic planners can discover a sequence of actions from initial to goal states given expert-defined, domain-specific logical action semantics. Large Language Models (LLMs) can di…