1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 1 cited
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
Chris Kelly, Luhui Hu, Jiayin Hu +7
The evolution of text to visual components facilitates people's daily lives, such as generating image, videos from text and identifying the desired elements within the images. Comp…
cs.CV2023
UnifiedVisionGPT: Streamlining Vision-Oriented AI through Generalized Multimodal Framework
Chris Kelly, Luhui Hu, Cindy Yang +6
In the current landscape of artificial intelligence, foundation models serve as the bedrock for advancements in both language and vision domains. OpenAI GPT-4 has emerged as the pi…