4 papers
Compression and Retrieval: Implicit Memory Retrieval for Video World Models
Zhan Peng, Jie Ma, Huiqiang Sun +6
Video world models hold promise for simulating interactive environments, yet maintaining consistent long-term memory across complex camera trajectories remains a critical challenge…
Robot Critics that Sweat the Small Stuff
Sruthi Sudhakar, Junbang Liang, Sreehari Rammohan +3
Large vision-language models contain several priors about the world and object interactions, making them useful critics during inference to steer robot policies towards success. Ho…
New York Smells: A Large Multimodal Dataset for Olfaction
Ege Ozguroglu, Junbang Liang, Ruoshi Liu +6
While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key bottleneck is the lack of divers…
Video Generators are Robot Policies
Junbang Liang, Pavel Tokmakov, Ruoshi Liu +4
Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or b…