collaborators

7 papers

cs.CV2026

RT-VLA: Real-Time Vision-Language-Action Models via Knowledge Distillation

Xiangyu Huang, Zhenlin Hua, Han Zhou +2

Vision-Language-Action (VLA) models have shown strong potential for end-to-end autonomous driving by jointly modeling visual perception, language reasoning, explainability and acti…

cs.CV2026

TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation

Hanyu Zhou, Chuanhao Ma, Gim Hee Lee

Vision-language-action (VLA) models perform well on training-seen robotic tasks but struggle to generalize to unseen scenes and objects. A key limitation lies in their implicit vis…

cs.CV2025

VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation

Hanyu Zhou, Chuanhao Ma, Gim Hee Lee

Vision-language-action (VLA) models show potential for general robotic tasks, but remain challenging in spatiotemporally coherent manipulation, which requires fine-grained represen…

cs.CV2025

Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation

Hanyu Zhou, Gim Hee Lee

Vision-language models (VLMs) have demonstrated strong performance in 2D scene understanding and generation, but extending this unification to the physical world remains an open ch…

cs.CV2025

STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene

Hanyu Zhou, Haonan Wang, Haoyue Liu +3

High-dynamic scene reconstruction aims to represent static background with rigid spatial features and dynamic objects with deformed continuous spatiotemporal features. Typically, e…

cs.CV2025

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding

Hanyu Zhou, Gim Hee Lee

Despite achieving significant progress in 2D image understanding, large multimodal models (LMMs) struggle in the physical world due to the lack of spatial representation. Typically…