From the 1 of 10 linked papers with an AI index.
9 papers
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
Zijun Lin, Zeqing Wang, Cheston Tan +2
The paper introduces StatePlay, a game world model that jointly predicts visual frames and internal game states using a mixture-of-transformers architecture to generate gameplay th…
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
Zijun Lin, Jiafei Duan, Haoquan Fang +4
Recent advances in robotic manipulation have integrated low-level robotic control into Vision-Language Models (VLMs), extending them into Vision-Language-Action (VLA) models. Altho…
SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
Zhaoxu Li, Chenqi Kong, Yi Yu +6
Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their re…
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
Shiqi Huang, Ziyue Wang, Zhongrong Zuo +3
Recent Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, existing RL pipelines rely he…
GEditBench v2: A Human-Aligned Benchmark for General Image Editing
Zhangqi Jiang, Zheng Sun, Xianfang Zeng +7
Recent advances in image editing have enabled models to handle complex instructions with impressive realism. However, existing evaluation frameworks lag behind: current benchmarks…
RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning
Shiqi Huang, Shuting He, Bihan Wen
Remote Sensing Visual Grounding (RSVG) aims to localize target objects in large-scale aerial imagery based on natural language descriptions. Owing to the vast spatial scale and hig…