10 papers
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications
Neel Mokaria, Rishie Raj, Dheeraj Baiju +13
Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestr…
Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
Wenxuan Zhang, Yuhui Wang, Donggang Jia +5
Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end t…
VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context
Xiaoqian Shen, Mohamed Elhoseiny
Large Vision Language Models (LVLMs) have achieved remarkable success on vision-language tasks, yet fine-grained perception over high-resolution images and long-context videos rema…
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions
Mohamed Elmoghany, Liangbing Zhao, Xiaoqian Shen +27
Generating long-form storytelling videos with consistent visual narratives remains a significant challenge in video synthesis. We present a novel framework, dataset, and a model th…
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
Xiaoqian Shen, Min-Hung Chen, Yu-Chiang Frank Wang +2
Grounded video question answering (GVQA) aims to localize relevant temporal segments in videos and generate accurate answers to a given question; however, large video-language mode…
iMotion-LLM: Instruction-Conditioned Trajectory Generation
Abdulwahab Felemban, Nussair Hroub, Jian Ding +4
We introduce iMotion-LLM, a large language model (LLM) integrated with trajectory prediction modules for interactive motion generation. Unlike conventional approaches, it generates…