8 papers · 1 filter
MVDGC: Joint 3D and 2D Multi-view Pedestrian Detection via Dual Geometric Constraints
Thinh Phan, Hao Vo, Khoa Vo +3
The core challenge in multi-view pedestrian detection (MVPD) lies in effective aggregation of visual features from different viewpoints for robust occlusion reasoning. Recent appro…
OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs
Hao Vo, Phu Loc Nguyen, Khoa Vo +7
Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Auton…
TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation
Duc Nguyen, Sieu Tran, Hao Vo +6
Unsupervised video object-centric learning aims to decompose dynamic scenes into temporally persistent entity representations. Existing recurrent video slot-attention methods propa…
ARGUSTRACK: A Multi-View Annotation System for Multi-Object Tracking
Hao Vo, Duc Nguyen, Ngan Le
Multi-Camera Multi-Target (MCMT) tracking has emerged as a critical capability for applications ranging from autonomous driving to animal behavior monitoring. While recent advances…
Dual-State Slot Attention: Decoupling Appearance and Identity for Video Object-Centric Learning
Sieu Tran, Duc Nguyen, Hao Vo +2
Unsupervised video object-centric learning aims to decompose dynamic scenes into persistent, object-level representations without supervision. However, existing slot-based methods…
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
Duy Le Dinh Anh, Kim Hoang Tran, Quang-Thuc Nguyen +1
Despite recent progress, Multi-Object Tracking (MOT) continues to face significant challenges, particularly its dependence on prior knowledge and predefined categories, complicatin…