collaborators

13 papers

cs.RO2026

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen +10

Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these videos is challenging: visually realistic ou…

cs.CV2026

StructSAM: Structure- and Spectrum-Preserving Token Merging for Segment Anything Models

Duy M. H. Nguyen, Tuan A. Tran, Duong Nguyen +17

Recent token merging techniques for Vision Transformers (ViTs) provide substantial speedups by reducing the number of tokens processed by self-attention, often without retraining.…

cs.CV2026

OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

Hao Vo, Phu Loc Nguyen, Khoa Vo +7

Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Auton…

cs.RO2026

Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think

Gia-Binh Nguyen, Trong-Bao Ho, Thien-Loc Ha +18

Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-billion parameter architectures impose pro…

cs.CV2026

TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation

Duc Nguyen, Sieu Tran, Hao Vo +6

Unsupervised video object-centric learning aims to decompose dynamic scenes into temporally persistent entity representations. Existing recurrent video slot-attention methods propa…

cs.CV2026

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

Duc Minh Nguyen, Nghiem Tuong Diep, Binh Gia Nguyen +20

Vision-Language-Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains…