9 citations · 9 across the 3 of their papers we have counts for
3 papers
cs.CV2025
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
Yang Zhou, Yifan Wang, Jianjun Zhou +16
The field of 4D world modeling - aiming to jointly capture spatial geometry and temporal dynamics - has witnessed remarkable progress in recent years, driven by advances in large-s…
cs.RO2023★ 9 cited
AlphaBlock: Embodied Finetuning for Vision-Language Reasoning in Robot Manipulation
Chuhao Jin, Wenhui Tan, Jiange Yang +4
We propose a novel framework for learning high-level cognitive capabilities in robot manipulation tasks, such as making a smiley face using building blocks. These tasks often invol…
cs.CV2023
CoMAE: Single Model Hybrid Pre-training on Small-Scale RGB-D Datasets
Jiange Yang, Sheng Guo, Gangshan Wu +1
Current RGB-D scene recognition approaches often train two standalone backbones for RGB and depth modalities with the same Places or ImageNet pre-training. However, the pre-trained…