5 papers · 1 filter
M3: High-fidelity Text-to-Image Generation via Multi-Modal, Multi-Agent and Multi-Round Visual Reasoning
Bangji Yang, Ruihan Guo, Jiajun Fan +2
Generative models have achieved impressive fidelity in text-to-image synthesis, yet struggle with complex compositional prompts involving multiple constraints. We introduce \textbf…
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
Chris Kelly, Luhui Hu, Jiayin Hu +7
The evolution of text to visual components facilitates people's daily lives, such as generating image, videos from text and identifying the desired elements within the images. Comp…
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
Chris Kelly, Luhui Hu, Bang Yang +7
With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achie…
WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
Deshun Yang, Luhui Hu, Yu Tian +5
Several text-to-video diffusion models have demonstrated commendable capabilities in synthesizing high-quality video content. However, it remains a formidable challenge pertaining…
UnifiedVisionGPT: Streamlining Vision-Oriented AI through Generalized Multimodal Framework
Chris Kelly, Luhui Hu, Cindy Yang +6
In the current landscape of artificial intelligence, foundation models serve as the bedrock for advancements in both language and vision domains. OpenAI GPT-4 has emerged as the pi…