3 papers
cs.CV2026
Visual Foresight for Robotic Stow: A Diffusion-Based World Model from Sparse Snapshots
Lijun Zhang, Nikhil Chacko, Petter Nilsson +8
Automated warehouses execute millions of stow operations, where robots place objects into storage bins. For these systems it is valuable to anticipate how a bin will look from the…
cs.CV2025
How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions
Aditya Prakash, Benjamin Lundell, Dmitry Andreychuk +3
We tackle the novel problem of predicting 3D hand motion and contact maps (or Interaction Trajectories) given a single RGB view, action text, and a 3D contact point on the object a…
cs.LG2025
Schema-Guided Scene-Graph Reasoning based on Multi-Agent Large Language Model System
Yiye Chen, Harpreet Sawhney, Nicholas Gydé +4
Scene graphs have emerged as a structured and serializable environment representation for grounded spatial reasoning with Large Language Models (LLMs). In this work, we propose SG^…