4.3k citations · 12k across the 32 of their papers we have counts for
4 papers · 1 filter
Visual Prompting via Image Inpainting
Amir Bar, Yossi Gandelsman, Trevor Darrell +2
How does one adapt a pre-trained visual model to novel downstream tasks without task-specific finetuning or any model modification? Inspired by prompting in NLP, this paper investi…
TL;DW? Summarizing Instructional Videos with Task Relevance & Cross-Modal Saliency
Medhini Narasimhan, Arsha Nagrani, Chen Sun +4
YouTube users looking for instructions for a specific task may spend a long time browsing content trying to find the right video that matches their needs. Creating a visual summary…
Disentangled Action Recognition with Knowledge Bases
Zhekun Luo, Shalini Ghosh, Devin Guillory +3
Action in video usually involves the interaction of human with objects. Action labels are typically composed of various combinations of verbs and nouns, but we may not have trainin…
Explaining Reinforcement Learning Policies through Counterfactual Trajectories
Julius Frost, Olivia Watkins, Eric Weiner +4
In order for humans to confidently decide where to employ RL agents for real-world tasks, a human developer must validate that the agent will perform well at test-time. Some policy…