collaborators

7 papers

cs.CV2026

MVDGC: Joint 3D and 2D Multi-view Pedestrian Detection via Dual Geometric Constraints

Thinh Phan, Hao Vo, Khoa Vo +3

The core challenge in multi-view pedestrian detection (MVPD) lies in effective aggregation of visual features from different viewpoints for robust occlusion reasoning. Recent appro…

cs.CV2026

OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

Hao Vo, Phu Loc Nguyen, Khoa Vo +7

Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Auton…

cs.RO2026

Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think

Gia-Binh Nguyen, Trong-Bao Ho, Thien-Loc Ha +18

Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-billion parameter architectures impose pro…

cs.CV2026

TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation

Duc Nguyen, Sieu Tran, Hao Vo +6

Unsupervised video object-centric learning aims to decompose dynamic scenes into temporally persistent entity representations. Existing recurrent video slot-attention methods propa…

cs.CV2026

ARGUSTRACK: A Multi-View Annotation System for Multi-Object Tracking

Hao Vo, Duc Nguyen, Ngan Le

Multi-Camera Multi-Target (MCMT) tracking has emerged as a critical capability for applications ranging from autonomous driving to animal behavior monitoring. While recent advances…

cs.RO2026

Self-Improving VLA Policies: Selected Diffusion Noise for Spurious-Robust Action Smoothing

Duc Minh Nguyen, Bao-Ngoc Dao, Tung M. Luu +15

Diffusion-based Vision-Language-Action (VLA) policies enable strong generalization in robotic manipulation, but remain sensitive to spurious visual correlations and noisy action ge…