activity
20232026
most citedUni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

2 citations · 3 across the 10 of their papers we have counts for

collaborators
Showing cs.ROShow all

11 papers · 1 filter

cs.RO2026

SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation

Songlin Wei, Zhenhao Ni, Jie Liu +9

Humanoid foundation models are advancing faster than we can evaluate them. While real-world testing is expensive and difficult to reproduce, existing simulation benchmarks focus pr…

cs.RO2026

: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation

Songlin Wei, Hongyi Jing, Boqian Li +12

We introduce (Psi-Zero), an open foundation model to address challenging humanoid loco-manipulation tasks. While existing approaches often attempt to address this fundamental…

cs.RO2026

ICLR: In-Context Imitation Learning with Visual Reasoning

Toan Nguyen, Weiduo Yuan, Songlin Wei +3

In-context imitation learning enables robots to adapt to new tasks from a small number of demonstrations without additional training. However, existing approaches typically conditi…

cs.RO2025

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

Shengliang Deng, Mi Yan, Songlin Wei +10

Embodied foundation models are gaining increasing attention for their zero-shot generalization, scalability, and adaptability to new tasks through few-shot post-training. However,…

cs.RO20251 cited

RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning

Haoran Geng, Feishi Wang, Songlin Wei +34

Data scaling and standardized evaluation benchmarks have driven significant advances in natural language processing and computer vision. However, robotics faces unique challenges i…

cs.RO20242 cited

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Jiazhao Zhang, Kunyu Wang, Shaoan Wang +6

A practical navigation agent must be capable of handling a wide range of interaction demands, such as following instructions, searching objects, answering questions, tracking peopl…