9 citations · 15 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 9 cited
Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
Rao Fu, Jingyu Liu, Xilun Chen +2
This paper introduces Scene-LLM, a 3D-visual-language model that enhances embodied agents' abilities in interactive 3D indoor environments by integrating the reasoning strengths of…
cs.CL2023★ 1 cited
The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task
Yifan Wu, Pengchuan Zhang, Wenhan Xiong +3
The study explores the effectiveness of the Chain-of-Thought approach, known for its proficiency in language tasks by breaking them down into sub-tasks and intermediate steps, in i…
cs.LG2023★ 5 cited
Jointly Training Large Autoregressive Multimodal Models
Emanuele Aiello, Lili Yu, Yixin Nie +2
In recent years, advances in the large-scale pretraining of language and text-to-image models have revolutionized the field of machine learning. Yet, integrating these two modaliti…