4 papers
Action Tokenizer Matters in In-Context Imitation Learning
An Dinh Vuong, Minh Nhat Vu, Dong An +1
In-context imitation learning (ICIL) is a new paradigm that enables robots to generalize from demonstrations to unseen tasks without retraining. A well-structured action representa…
Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder
An Dinh Vuong, Minh Nhat Vu, Ian Reid
Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities. Recent geometry-…
HabiCrowd: A High Performance Simulator for Crowd-Aware Visual Navigation
An Dinh Vuong, Toan Tien Nguyen, Minh Nhat VU +5
Visual navigation, a foundational aspect of Embodied AI (E-AI), has been significantly studied in the past few years. While many 3D simulators have been introduced to support visua…
Language-driven Grasp Detection
An Dinh Vuong, Minh Nhat Vu, Baoru Huang +4
Grasp detection is a persistent and intricate challenge with various industrial applications. Recently, many methods and datasets have been proposed to tackle the grasp detection p…