5 papers
Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations
Zichao Li, Gang Wu, Zichao Wang +5
Large language model agents operate in partially observable, long-horizon settings where obtaining supervision remains a major bottleneck. We address this by utilizing a source of…
Anticipatory Planning for Multimodal AI Agents
Yongyuan Liang, Shijie Zhou, Yu Gu +6
Recent advances in multimodal agents have improved computer-use interaction and tool-usage, yet most existing systems remain reactive, optimizing actions in isolation without reaso…
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions
Mohamed Elmoghany, Liangbing Zhao, Xiaoqian Shen +27
Generating long-form storytelling videos with consistent visual narratives remains a significant challenge in video synthesis. We present a novel framework, dataset, and a model th…
Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
Kunjal Panchal, Saayan Mitra, Somdeb Sarkhel +4
Recent advances in video-language models have enabled powerful applications like video retrieval, captioning, and assembly. However, executing such multi-stage pipelines efficientl…
GUI Agents: A Survey
Dang Nguyen, Jian Chen, Yu Wang +27
Graphical User Interface (GUI) agents, powered by Large Foundation Models, have emerged as a transformative approach to automating human-computer interaction. These agents autonomo…