7 papers
Pelican-VLA 0.5: Attending Before Acting Benefits Generalization
Zeyuan Ding, Wenhai Liu, Yang Xu +6
In this report, we present Pelican-VLA 0.5, a unified VLA model that integrates vision-language understanding, future-frame generation, and action prediction within a single archit…
VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing
Haoyuan Shi, Xiancong Ren, Yingji Zhang +9
Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Trace, a progressive diagnostic…
Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action
Yi Zhang, Yinda Chen, Che Liu +26
We present Pelican-Unify 1.0, the first embodied foundation model trained according to the principle of unification. Pelican-Unify 1.0 uses a single VLM as a unified understanding…
EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
Haozhe Shan, Xiancong Ren, Han Dong +9
While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on question-answering or multiple-…
A2Eval: Agentic and Automated Evaluation for Embodied Brain
Shuai Zhang, Jiayu Hu, Zijie Chen +9
Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor intensive paradigm…
Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization
Yi Zhang, Che Liu, Xiancong Ren +17
Developing a universal and versatile embodied intelligence system presents two primary challenges: the critical embodied data bottleneck, where real-world data is scarce and expens…