21 papers
GRAIN: Group Aggregation via Min-Norm Objective
Nghia Bui, Jiarui Yao, Lijing Wang
Learning instability is a long-standing problem across machine learning, but it is especially acute in the overparameterized regime that defines modern deep learning: large models…
OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs
Hao Vo, Phu Loc Nguyen, Khoa Vo +7
Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Auton…
Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think
Gia-Binh Nguyen, Trong-Bao Ho, Thien-Loc Ha +18
Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-billion parameter architectures impose pro…
TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation
Duc Nguyen, Sieu Tran, Hao Vo +6
Unsupervised video object-centric learning aims to decompose dynamic scenes into temporally persistent entity representations. Existing recurrent video slot-attention methods propa…
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
Hao Vo, Khoa Vo, Phu Loc Nguyen +10
Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity acros…
Documentation-Guided Agentic Codebase Migration from C to Rust
Minh Le-Anh, Anh Nguyen Hoang, Bach Le +1
Migrating legacy C repositories to Rust promises stronger memory safety, but existing translators often work at the level of files or functions and miss architectural intent. We pr…