2 papers
cs.CV2026
Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers
Cong Liu, Xiaofang Li, Simon X. Yang
Vision Transformers (ViTs) commonly rely on injected positional mechanisms to address self-attention's permutation invariance. Motivated by the spatial regularities of natural imag…
cs.RO2026
MemoryWAM: Efficient World Action Modeling with Persistent Memory
Sizhe Yang, Juncheng Mu, Tianming Wei +8
Robust robotic manipulation in the real world requires not only an understanding of the current observation, but also memory and dynamics modeling. World action models (WAMs) posse…