2 papers
cs.CV2024
Mora: Enabling Generalist Video Generation via A Multi-Agent Framework
Zhengqing Yuan, Yixin Liu, Yihan Cao +10
Text-to-video generation has made significant strides, but replicating the capabilities of advanced systems like OpenAI Sora remains challenging due to their closed-source nature.…
cs.CV2024
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
Zhengqing Yuan, Zhaoxu Li, Weiran Huang +2
In recent years, multimodal large language models (MLLMs) such as GPT-4V have demonstrated remarkable advancements, excelling in a variety of vision-language tasks. Despite their p…