7 papers
RT-VLA: Real-Time Vision-Language-Action Models via Knowledge Distillation
Xiangyu Huang, Zhenlin Hua, Han Zhou +2
Vision-Language-Action (VLA) models have shown strong potential for end-to-end autonomous driving by jointly modeling visual perception, language reasoning, explainability and acti…
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
Hanyu Zhou, Chuanhao Ma, Gim Hee Lee
Vision-language-action (VLA) models perform well on training-seen robotic tasks but struggle to generalize to unseen scenes and objects. A key limitation lies in their implicit vis…
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
Hanyu Zhou, Chuanhao Ma, Gim Hee Lee
Vision-language-action (VLA) models show potential for general robotic tasks, but remain challenging in spatiotemporally coherent manipulation, which requires fine-grained represen…
Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
Hanyu Zhou, Gim Hee Lee
Vision-language models (VLMs) have demonstrated strong performance in 2D scene understanding and generation, but extending this unification to the physical world remains an open ch…
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
Hanyu Zhou, Haonan Wang, Haoyue Liu +3
High-dynamic scene reconstruction aims to represent static background with rigid spatial features and dynamic objects with deformed continuous spatiotemporal features. Typically, e…
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
Hanyu Zhou, Gim Hee Lee
Despite achieving significant progress in 2D image understanding, large multimodal models (LMMs) struggle in the physical world due to the lack of spatial representation. Typically…