3 papers
cs.CV2026
Probing Visual Planning in Image Editing Models
Zhimu Zhou, Yanpeng Zhao, Qiuyu Liao +2
Visual planning represents a crucial facet of human intelligence, especially in tasks that require complex spatial reasoning and navigation. Yet, in machine learning, this inherent…
cs.CV2025
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
Chenghao Liu, Jiachen Zhang, Chengxuan Li +4
Vision-Language-Action (VLA) models process visual inputs independently at each timestep, discarding valuable temporal information inherent in robotic manipulation tasks. This fram…
cs.CV2025
MSNav: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and LLM Spatial Reasoning
Chenghao Liu, Zhimu Zhou, Jiachen Zhang +3
Vision-and-Language Navigation (VLN) requires an agent to interpret natural language instructions and navigate complex environments. Current approaches often adopt a "black-box" pa…