From the 1 of 9 linked papers with an AI index.
9 papers
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
Zihan Dong, Rui Qian, Qishi Zhan +3
The paper introduces Adaptive Anticipatory Policy Trees (AAPT), a method that pre‑computes conditional action trees during idle screen time so GUI agents can react instantly to eve…
Scaling GUI Agents with Visual State Transitions
Xiangyan Liu, Kaixin Li, Haonan Wang +6
We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state trans…
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
Yunlong Lin, Zixu Lin, Zhaohu Xing +23
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, aud…
From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms
Jinghao Luo, Yuchen Tian, Chuxue Cao +6
Large Language Model (LLM)-based agents have fundamentally reshaped artificial intelligence by integrating external tools and planning capabilities. While memory mechanisms have em…
MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique
Gailun Zeng, Ziyang Luo, Hongzhan Lin +5
The ability of critique is vital for models to self-improve and serve as reliable AI assistants. While extensively studied in language-only settings, multimodal critique of Large M…
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
Zixin Chen, Hongzhan Lin, Kaixin Li +3
The proliferation of memes on social media necessitates the capabilities of multimodal Large Language Models (mLLMs) to effectively understand multimodal harmfulness. Existing eval…