2 papers
cs.GR2026
CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation
Juyeong Hwang, Seong-Eun Hong, Jinhyun Kim +4
Crowds do not merely move; they decide. Human navigation is inherently contextual: people interpret the meaning of space, social norms, and potential consequences before acting. Si…
cs.CV2026
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
Jisoo Kim, Jungbin Cho, Sanghyeok Chu +9
Humans learn not only how their bodies move, but also how the surrounding world responds to their actions. In contrast, while recent Vision-Language-Action (VLA) models exhibit imp…