3 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
Chris Kelly, Luhui Hu, Jiayin Hu +7
The evolution of text to visual components facilitates people's daily lives, such as generating image, videos from text and identifying the desired elements within the images. Comp…
cs.CV2024★ 3 cited
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
Chris Kelly, Luhui Hu, Bang Yang +7
With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achie…
cs.CV2023
UnifiedVisionGPT: Streamlining Vision-Oriented AI through Generalized Multimodal Framework
Chris Kelly, Luhui Hu, Cindy Yang +6
In the current landscape of artificial intelligence, foundation models serve as the bedrock for advancements in both language and vision domains. OpenAI GPT-4 has emerged as the pi…