From the 1 of 141 linked papers with an AI index.
10 papers · 1 filter
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications
Neel Mokaria, Rishie Raj, Dheeraj Baiju +13
Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestr…
ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents
Binjie Zhang, Mike Zheng Shou
Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Existing works have two common g…
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
Meng Chu, Xuan Billy Zhang, Kevin Qinghong Lin +47
As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a central bottleneck. Agents that ma…
WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point
Henry Hengyuan Zhao, Kaiming Yang, Wendi Yu +2
Recent progress in GUI agents has substantially improved visual grounding, yet robust planning remains challenging, particularly when the environment deviates from a canonical init…
AUTO-Explorer: Automated Data Collection for GUI Agent
Xiangwu Guo, Difei Gao, Mike Zheng Shou
Recent advancements in GUI agents have significantly expanded their ability to interpret natural language commands to manage software interfaces. However, acquiring GUI data remain…
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
Jiaqi Wang, Kevin Qinghong Lin, James Cheng +1
Reinforcement Learning (RL) has proven to be an effective post-training strategy for enhancing reasoning in vision-language models (VLMs). Group Relative Policy Optimization (GRPO)…