12 papers
ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation
Haonan Wang, Hanyu Zhou, Tao Gu +1
Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spatiotemporal scale. Typically,…
ST-: Structured SpatioTemporal VLA for Robotic Manipulation
Chuanhao Ma, Hanyu Zhou, Shihan Peng +3
Vision-language-action (VLA) models have achieved great success on general robotic tasks, but still face challenges in fine-grained spatiotemporal manipulation. Typically, existing…
NEC-Diff: Noise-Robust Event-RAW Complementary Diffusion for Seeing Motion in Extreme Darkness
Haoyue Liu, Jinghan Xu, Luxin Feng +4
High-quality imaging of dynamic scenes in extremely low-light conditions is highly challenging. Photon scarcity induces severe noise and texture loss, causing significant image deg…
Cog2Gen3D: Sculpturing 3D Semantic-Geometric Cognition for 3D Generation
Haonan Wang, Hanyu Zhou, Haoyue Liu +2
Generative models have achieved success in producing semantically plausible 2D images, but it remains challenging in 3D generation due to the absence of spatial geometry constraint…
Adapting Depth Anything to Adverse Imaging Conditions with Events
Shihan Peng, Yuyang Xiong, Hanyu Zhou +5
Robust depth estimation under dynamic and adverse lighting conditions is essential for robotic systems. Currently, depth foundation models, such as Depth Anything, achieve great su…
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
Haonan Wang, Hanyu Zhou, Haoyue Liu +1
We investigate a challenging task of dynamic scene geometry estimation, which requires representing both spatial and temporal features. Typically, existing methods align the two fe…