40 citations · 49 across the 3 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023★ 40 cited
3D-LLM: Injecting the 3D World into Large Language Models
Yining Hong, Haoyu Zhen, Peihao Chen +4
Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are…
cs.CV2023★ 8 cited
See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning
Zhenfang Chen, Qinhong Zhou, Yikang Shen +3
Large pre-trained vision and language models have demonstrated remarkable capacities for various tasks. However, solving the knowledge-based visual reasoning tasks remains challeng…