Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation
Yongxin Wang, Meng Cao, Haokun Lin +5
Multimodal large language models (MLLMs) have achieved remarkable progress on various visual question answering and reasoning tasks leveraging instruction fine-tuning specific data…
cs.CV2024
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration
Panwen Hu, Jin Jiang, Jianqi Chen +4
The advent of AI-Generated Content (AIGC) has spurred research into automated video generation to streamline conventional processes. However, automating storytelling video producti…