5 papers · 1 filter
Trimming the Long-Tail of Visual World Modeling Evaluation
Bingxuan Li, Yining Hong, Cheng Qian +6
Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human experience and visual data, while a broad spectrum of rare and irr…
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
Yining Hong, Jiageng Liu, Han Yin +5
Spatial intelligence unfolds through a perception-action loop: agents act to acquire observations, and reason about how observations vary as a function of action. Rather than passi…
Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
Yuming Gu, Yizhi Wang, Yining Hong +9
Embodied visual planning aims to enable manipulation tasks by imagining how a scene evolves toward a desired goal and using the imagined trajectories to guide actions. Video diffus…
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
Wenbo Hu, Yining Hong, Yanjun Wang +7
Humans excel at performing complex tasks by leveraging long-term memory across temporal and spatial experiences. In contrast, current Large Language Models (LLMs) struggle to effec…
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
Yining Hong, Beide Liu, Maxine Wu +9
Human beings are endowed with a complementary learning system, which bridges the slow learning of general world dynamics with fast storage of episodic memory from a new experience.…