40 citations · 57 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 14 cited
3D-VLA: A 3D Vision-Language-Action Generative World Model
Haoyu Zhen, Xiaowen Qiu, Peihao Chen +5
Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by le…
cs.CV2023★ 3 cited
CHORD: Category-level Hand-held Object Reconstruction via Shape Deformation
Kailin Li, Lixin Yang, Haoyu Zhen +6
In daily life, humans utilize hands to manipulate objects. Modeling the shape of objects that are manipulated by the hand is essential for AI to comprehend daily tasks and to learn…
cs.CV2023★ 40 cited
3D-LLM: Injecting the 3D World into Large Language Models
Yining Hong, Haoyu Zhen, Peihao Chen +4
Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are…