most citedWorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs

4 citations · 9 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20241 cited

VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding

Chris Kelly, Luhui Hu, Jiayin Hu +7

The evolution of text to visual components facilitates people's daily lives, such as generating image, videos from text and identifying the desired elements within the images. Comp…

cs.CV20243 cited

VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework

Chris Kelly, Luhui Hu, Bang Yang +7

With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achie…

cs.CV20244 cited

WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs

Deshun Yang, Luhui Hu, Yu Tian +5

Several text-to-video diffusion models have demonstrated commendable capabilities in synthesizing high-quality video content. However, it remains a formidable challenge pertaining…

cs.CV20241 cited

Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning

Bang Yang, Yong Dai, Xuxin Cheng +3

While vision-language pre-trained models (VL-PTMs) have advanced multimodal research in recent years, their mastery in a few languages like English restricts their applicability in…

cs.CV2023

UnifiedVisionGPT: Streamlining Vision-Oriented AI through Generalized Multimodal Framework

Chris Kelly, Luhui Hu, Cindy Yang +6

In the current landscape of artificial intelligence, foundation models serve as the bedrock for advancements in both language and vision domains. OpenAI GPT-4 has emerged as the pi…