4 papers
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
Yunxin Li, Zhenyu Liu, Zitao Li +19
Reasoning lies at the heart of intelligence, shaping the ability to make decisions, draw conclusions, and generalize across domains. In artificial intelligence, as systems increasi…
A Unified Agentic Framework for Evaluating Conditional Image Generation
Jifang Wang, Xue Yang, Longyue Wang +7
Conditional image generation has gained significant attention for its ability to personalize content. However, the field faces challenges in developing task-agnostic, reliable, and…
FilmAgent: A Multi-Agent Framework for End-to-End Film Automation in Virtual 3D Spaces
Zhenran Xu, Longyue Wang, Jifang Wang +7
Virtual film production requires intricate decision-making processes, including scriptwriting, virtual cinematography, and precise actor positioning and actions. Motivated by recen…
Medico: Towards Hallucination Detection and Correction with Multi-source Evidence Fusion
Xinping Zhao, Jindi Yu, Zhenyu Liu +5
As we all know, hallucinations prevail in Large Language Models (LLMs), where the generated content is coherent but factually incorrect, which inflicts a heavy blow on the widespre…