collaborators

10 papers

cs.AI2026

A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications

Neel Mokaria, Rishie Raj, Dheeraj Baiju +13

Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestr…

cs.AI2026

Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning

Wenxuan Zhang, Yuhui Wang, Donggang Jia +5

Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end t…

cs.CV2026

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context

Xiaoqian Shen, Mohamed Elhoseiny

Large Vision Language Models (LVLMs) have achieved remarkable success on vision-language tasks, yet fine-grained perception over high-resolution images and long-context videos rema…

cs.CV2026

InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions

Mohamed Elmoghany, Liangbing Zhao, Xiaoqian Shen +27

Generating long-form storytelling videos with consistent visual narratives remains a significant challenge in video synthesis. We present a novel framework, dataset, and a model th…

cs.CV2025

Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in

Xiaoqian Shen, Min-Hung Chen, Yu-Chiang Frank Wang +2

Grounded video question answering (GVQA) aims to localize relevant temporal segments in videos and generate accurate answers to a given question; however, large video-language mode…

cs.CV2025

iMotion-LLM: Instruction-Conditioned Trajectory Generation

Abdulwahab Felemban, Nussair Hroub, Jian Ding +4

We introduce iMotion-LLM, a large language model (LLM) integrated with trajectory prediction modules for interactive motion generation. Unlike conventional approaches, it generates…