1 citations · 1 across the 5 of their papers we have counts for
6 papers
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
Yidong Huang, Zun Wang, Han Lin +5
Recent video generation approaches increasingly rely on planning intermediate control signals such as object trajectories to improve temporal coherence and motion fidelity. However…
Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
Jason Jabbour, Dong-Ki Kim, Max Smith +6
Vision-Language-Action (VLA) models have advanced robotic capabilities but remain challenging to deploy on resource-limited hardware. Pruning has enabled efficient compression of l…
VENTURA: Adapting Image Diffusion Models for Unified Task Conditioned Navigation
Arthur Zhang, Xiangyun Meng, Luca Calliari +5
Robots must adapt to diverse human instructions and operate safely in unstructured, open-world environments. Recent Vision-Language models (VLMs) offer strong priors for grounding…
StageACT: Stage-Conditioned Imitation for Robust Humanoid Door Opening
Moonyoung Lee, Dong Ki Kim, Jai Krishna Bandi +4
Humanoid robots promise to operate in everyday human environments without requiring modifications to the surroundings. Among the many skills needed, opening doors is essential, as…
Enter the Mind Palace: Reasoning and Planning for Long-term Active Embodied Question Answering
Muhammad Fadhil Ginting, Dong-Ki Kim, Xiangyun Meng +10
As robots become increasingly capable of operating over extended periods -- spanning days, weeks, and even months -- they are expected to accumulate knowledge of their environments…
SayComply: Grounding Field Robotic Tasks in Operational Compliance through Retrieval-Based Language Models
Muhammad Fadhil Ginting, Dong-Ki Kim, Sung-Kyun Kim +4
This paper addresses the problem of task planning for robots that must comply with operational manuals in real-world settings. Task planning under these constraints is essential fo…