collaborators

6 papers

cs.RO2026

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen +10

Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these videos is challenging: visually realistic ou…

cs.RO2026

Self-Improving VLA Policies: Selected Diffusion Noise for Spurious-Robust Action Smoothing

Duc Minh Nguyen, Bao-Ngoc Dao, Tung M. Luu +15

Diffusion-based Vision-Language-Action (VLA) policies enable strong generalization in robotic manipulation, but remain sensitive to spurious visual correlations and noisy action ge…

cs.CV2026

Enhancing Few-Shot Classification of Benchmark and Disaster Imagery with ABHFA-Net

Gao Yu Lee, Tanmoy Dam, Md Meftahul Ferdaus +2

The rising incidence of natural and human-induced disasters necessitates robust visual recognition systems capable of operating under limited labeled data conditions. However, disa…

cs.CL2026

Safety-Oriented Evaluation of Language Understanding Systems for Air Traffic Control

Yujing Chang, Yash Guleria, Duc-Thinh Pham +4

Air Traffic Control (ATC) is a safety-critical domain in which incorrect interpretation of instructions may lead to severe operational consequences. While large language models (LL…

cs.CV2025

ANROT-HELANet: Adverserially and Naturally Robust Attention-Based Aggregation Network via The Hellinger Distance for Few-Shot Classification

Gao Yu Lee, Tanmoy Dam, Md Meftahul Ferdaus +2

Few-Shot Learning (FSL), which involves learning to generalize using only a few data samples, has demonstrated promising and superior performances to ordinary CNN methods. While Ba…

cs.CV2025

DRACO-DehazeNet: An Efficient Image Dehazing Network Combining Detail Recovery and a Novel Contrastive Learning Paradigm

Gao Yu Lee, Tanmoy Dam, Md Meftahul Ferdaus +2

Image dehazing is crucial for clarifying images obscured by haze or fog, but current learning-based approaches is dependent on large volumes of training data and hence consumed sig…