activity
20242026
collaborators

7 papers

cs.RO2026

Robot Critics that Sweat the Small Stuff

Sruthi Sudhakar, Junbang Liang, Sreehari Rammohan +3

Large vision-language models contain several priors about the world and object interactions, making them useful critics during inference to steer robot policies towards success. Ho…

cs.CV2026

: Smaller Self-Supervised ViTs Localize Better than Larger Ones

Sreehari Rammohan, Huy Ha, Carl Vondrick

Robust visual classification often depends on localizing the main foreground objects in an image while ignoring contextual distractors. Surprisingly, we find that the attention map…

cs.CV2025

New York Smells: A Large Multimodal Dataset for Olfaction

Ege Ozguroglu, Junbang Liang, Ruoshi Liu +6

While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key bottleneck is the lack of divers…

cs.RO2025

Video Generators are Robot Policies

Junbang Liang, Pavel Tokmakov, Ruoshi Liu +4

Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or b…

cs.CV2024

EraseDraw: Learning to Draw Step-by-Step via Erasing Objects from Images

Alper Canberk, Maksym Bondarenko, Ege Ozguroglu +2

Creative processes such as painting often involve creating different components of an image one by one. Can we build a computational model to perform this task? Prior works often f…

cs.RO2024

Self-Improving Autonomous Underwater Manipulation

Ruoshi Liu, Huy Ha, Mengxue Hou +2

Underwater robotic manipulation faces significant challenges due to complex fluid dynamics and unstructured environments, causing most manipulation systems to rely heavily on human…