9 papers
Threading Optimization for Vision-Language-Action Model Inference in Low-Cost Smart Agricultural Manipulation
Keith Truongcao, Christopher Nhu, Zijian An +3
Vision-Language Action (VLA) models continue to face challenges such as slow inference speed and difficulty performing fine-grained motion adjustments, limiting their widespread ad…
ROG-Grasp: Root-Oriented Geometry for Robotic Grasping and Placement
Zijian An, Augustus Sroka, Ran Yang +8
Orientation-aware manipulation is essential in post-harvest agricultural processing, where produce must be grasped and placed in consistent configurations. This paper presents ROG-…
CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping
Zijian An, Ran Yang, Yiming Feng +1
Vision-language-action (VLA) models have recently emerged as a promising paradigm for robotic control, enabling end-to-end policies that ground natural language instructions into v…
VILAS: A VLA-Integrated Low-cost Architecture with Soft Grasping for Robotic Manipulation
Zijian An, Hadi Khezam, Bill Cai +5
We present VILAS, a fully low-cost, modular robotic manipulation platform designed to support end-to-end vision-language-action (VLA) policy learning and deployment on accessible h…
Large Language Models for Multi-Robot Systems: A Survey
Peihan Li, Zijian An, Shams Abrar +1
The rapid advancement of Large Language Models (LLMs) has opened new possibilities in Multi-Robot Systems (MRS), enabling enhanced communication, task allocation and planning, and…
Vision Language Models Cannot Plan, but Can They Formalize?
Muyu He, Yuxi Zheng, Yuchen Liu +7
The advancement of vision language models (VLMs) has empowered embodied agents to accomplish simple multimodal planning tasks, but not long-horizon ones requiring long sequences of…