From the 1 of 10 linked papers with an AI index.
10 papers
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
Zijun Lin, Zeqing Wang, Cheston Tan +2
The paper introduces StatePlay, a game world model that jointly predicts visual frames and internal game states using a mixture-of-transformers architecture to generate gameplay th…
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
Zijun Lin, Jiafei Duan, Haoquan Fang +4
Recent advances in robotic manipulation have integrated low-level robotic control into Vision-Language Models (VLMs), extending them into Vision-Language-Action (VLA) models. Altho…
Information Fidelity in Tool-Using LLM Agents: A Martingale Analysis of the Model Context Protocol
Flint Xiaofeng Fan, Cheston Tan, Roger Wattenhofer +1
As AI agents powered by large language models (LLMs) increasingly use external tools for high-stakes decisions, a critical reliability question arises: how do errors propagate acro…
TangramSR: Can Vision-Language Models Reason in Continuous Geometric Space?
Yikun Zong, Cheston Tan
Humans excel at spatial reasoning tasks like Tangram puzzle assembly through cognitive processes involving mental rotation, iterative refinement, and visual feedback. Inspired by h…
From Grunts to Lexicons: Emergent Language from Cooperative Foraging
Maytus Piriyajitakonkij, Rujikorn Charakorn, Weicheng Tao +4
Language is a powerful communicative and cognitive tool. It enables humans to express thoughts, share intentions, and reason about complex phenomena. Despite our fluency in using a…
Stencil: Subject-Driven Generation with Context Guidance
Gordon Chen, Ziqi Huang, Cheston Tan +1
Recent text-to-image diffusion models can generate striking visuals from text prompts, but they often fail to maintain subject consistency across generations and contexts. One majo…