Publications (153)
Egocentric Video Task Translation @ Ego4D Challenge 2022
Zihui Xue, Yale Song, Kristen Grauman +1
PONI: Potential Functions for ObjectGoal Navigation with Interaction-free Learning
Santhosh Kumar Ramakrishnan, Devendra Singh Chaplot, Ziad Al-Halah +2
Learning image representations tied to ego-motion
Dinesh Jayaraman, Kristen Grauman
Learning Affordance Landscapes for Interaction Exploration in 3D Environments
Tushar Nagarajan, Kristen Grauman
Ego4D: Around the World in 3,000 Hours of Egocentric Video
Kristen Grauman, Andrew Westbury, Eugene Byrne +82
Visual Acoustic Matching
Changan Chen, Ruohan Gao, Paul Calamia +1
Retrospectives on the Embodied AI Workshop
Matt Deitke, Dhruv Batra, Yonatan Bisk +36
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
Changan Chen, Kumar Ashutosh, Rohit Girdhar +2
EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding
Shuhan Tan, Tushar Nagarajan, Kristen Grauman
ShapeCodes: Self-Supervised Feature Learning by Lifting Views to Viewgrids
Dinesh Jayaraman, Ruohan Gao, Kristen Grauman
Co-Separating Sounds of Visual Objects
Ruohan Gao, Kristen Grauman
Seeing without Pixels: Perception from Camera Trajectories
Zihui Xue, Kristen Grauman, Dima Damen +2
Novel-View Acoustic Synthesis
Changan Chen, Alexander Richard, Roman Shapovalov +4
Single-Stage Visual Query Localization in Egocentric Videos
Hanwen Jiang, Santhosh Kumar Ramakrishnan, Kristen Grauman
When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
Mi Luo, Zihui Xue, Alex Dimakis +1
Grounded Human-Object Interaction Hotspots from Video (Extended Abstract)
Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman
Learning Skill-Attributes for Transferable Assessment in Video
Kumar Ashutosh, Kristen Grauman
Semantic Jitter: Dense Supervision for Visual Comparisons via Synthetic Images
Aron Yu, Kristen Grauman
Active Audio-Visual Separation of Dynamic Sound Sources
Sagnik Majumder, Kristen Grauman
SoundSpaces: Audio-Visual Navigation in 3D Environments
Changan Chen, Unnat Jain, Carl Schissler +5
Look-ahead before you leap: end-to-end active recognition by forecasting the effect of motion
Dinesh Jayaraman, Kristen Grauman
Seeing the Arrow of Time in Large Multimodal Models
Zihui Xue, Mi Luo, Kristen Grauman
Audio-Visual Floorplan Reconstruction
Senthil Purushwalkam, Sebastian Vicenc Amengual Gari, Vamsi Krishna Ithapu +4
Sidekick Policy Learning for Active Visual Exploration
Santhosh K. Ramakrishnan, Kristen Grauman
Large-Margin Determinantal Point Processes
Boqing Gong, Wei-lun Chao, Kristen Grauman +1
HierVL: Learning Hierarchical Video-Language Embeddings
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani +1
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
Ami Baid, Zihui Xue, Kristen Grauman
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
Zihui Xue, Mi Luo, Changan Chen +1
Compare and Contrast: Learning Prominent Visual Differences
Steven Chen, Kristen Grauman
Chat2Map: Efficient Scene Mapping from Multi-Ego Conversations
Sagnik Majumder, Hao Jiang, Pierre Moulon +4
Predicting Foreground Object Ambiguity and Efficiently Crowdsourcing the Segmentation(s)
Danna Gurari, Kun He, Bo Xiong +6
WhittleSearch: Interactive Image Search with Relative Attribute Feedback
Adriana Kovashka, Devi Parikh, Kristen Grauman
Human detectors are surprisingly powerful reward models
Kumar Ashutosh, XuDong Wang, Xi Yin +4
ViBE: Dressing for Diverse Body Shapes
Wei-Lin Hsiao, Kristen Grauman
BlockDrop: Dynamic Inference Paths in Residual Networks
Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar +4
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi +26
Listen to Look: Action Recognition by Previewing Audio
Ruohan Gao, Tae-Hyun Oh, Kristen Grauman +1
Object-Centric Representation Learning from Unlabeled Videos
Ruohan Gao, Dinesh Jayaraman, Kristen Grauman
Pano2Vid: Automatic Cinematography for Watching 360$^{\circ}$ Videos
Yu-Chuan Su, Dinesh Jayaraman, Kristen Grauman
What You Say Is What You Show: Visual Narration Detection in Instructional Videos
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani +1
Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos
Mi Luo, Zihui Xue, Alex Dimakis +1
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98
Shapes as Product Differentiation: Neural Network Embedding in the Analysis of Markets for Fonts
Sukjin Han, Eric H. Schulman, Kristen Grauman +1
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
Joungbin An, Agrim Jain, Kristen Grauman
VisualEchoes: Spatial Image Representation Learning through Echolocation
Ruohan Gao, Changan Chen, Ziad Al-Halah +2
Efficient Activity Detection in Untrimmed Video with Max-Subgraph Search
Chao-Yeh Chen, Kristen Grauman
Snap Angle Prediction for 360$^{\circ}$ Panoramas
Bo Xiong, Kristen Grauman
From Culture to Clothing: Discovering the World Events Behind A Century of Fashion Images
Wei-Lin Hsiao, Kristen Grauman
Personal Visual Context Learning in Large Multimodal Models
Zihui Xue, Ami Baid, Sangho Kim +2
Progress-Aware Video Frame Captioning
Zihui Xue, Joungbin An, Xitong Yang +1
Video Analysis for Body-worn Cameras in Law Enforcement
Jason J. Corso, Alexandre Alahi, Kristen Grauman +4
Learning Compressible 360° Video Isomers
Yu-Chuan Su, Kristen Grauman
Im2Flow: Motion Hallucination from Static Images for Action Recognition
Ruohan Gao, Bo Xiong, Kristen Grauman
NaQ: Leveraging Narrations as Queries to Supervise Episodic Memory
Santhosh Kumar Ramakrishnan, Ziad Al-Halah, Kristen Grauman
Discovering Underground Maps from Fashion
Utkarsh Mall, Kavita Bala, Tamara Berg +1
Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos
Yanghao Li, Tushar Nagarajan, Bo Xiong +1
FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos
Suyog Dutt Jain, Bo Xiong, Kristen Grauman
DexVIP: Learning Dexterous Grasping with Human Hand Pose Priors from Video
Priyanka Mandikal, Kristen Grauman
Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback
Hui Wu, Yupeng Gao, Xiaoxiao Guo +4
SpotTune: Transfer Learning through Adaptive Fine-tuning
Yunhui Guo, Honghui Shi, Abhishek Kumar +3
Emergence of Exploratory Look-Around Behaviors through Active Observation Completion
Santhosh K. Ramakrishnan, Dinesh Jayaraman, Kristen Grauman
Learning the Latent "Look": Unsupervised Discovery of a Style-Coherent Embedding from Fashion Images
Wei-Lin Hsiao, Kristen Grauman
Learning Audio-Visual Dereverberation
Changan Chen, Wei Sun, David Harwath +1
An Exploration of Embodied Visual Exploration
Santhosh K. Ramakrishnan, Dinesh Jayaraman, Kristen Grauman
You2Me: Inferring Body Pose in Egocentric Video via First and Second Person Interactions
Evonne Ng, Donglai Xiang, Hanbyul Joo +1
Predicting Important Objects for Egocentric Video Summarization
Yong Jae Lee, Kristen Grauman
Occupancy Anticipation for Efficient Exploration and Navigation
Santhosh K. Ramakrishnan, Ziad Al-Halah, Kristen Grauman
Detangling People: Individuating Multiple Close People and Their Body Parts via Region Assembly
Hao Jiang, Kristen Grauman
Leaving Some Stones Unturned: Dynamic Feature Prioritization for Activity Detection in Streaming Video
Yu-Chuan Su, Kristen Grauman
EGO-TOPO: Environment Affordances from Egocentric Video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer +1
VizWiz Grand Challenge: Answering Visual Questions from Blind People
Danna Gurari, Qing Li, Abigale J. Stangl +5
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah +1
Grounded Human-Object Interaction Hotspots from Video
Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman
Extreme Relative Pose Estimation for RGB-D Scans via Scene Completion
Zhenpei Yang, Jeffrey Z. Pan, Linjie Luo +3
ExpertAF: Expert Actionable Feedback from Video
Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos +2
Egocentric Activity Recognition and Localization on a 3D Map
Miao Liu, Lingni Ma, Kiran Somasundaram +4
Few-Shot Audio-Visual Learning of Environment Acoustics
Sagnik Majumder, Changan Chen, Ziad Al-Halah +1
A Domain-Agnostic Approach for Characterization of Lifelong Learning Systems
Megan M. Baker, Alexander New, Mario Aguilar-Simon +44
Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias
Krishna Kumar Singh, Dhruv Mahajan, Kristen Grauman +3
Mash, Spread, Slice! Learning to Manipulate Object States via Visual Spatial Progress
Priyanka Mandikal, Jiaheng Hu, Shivin Dass +3
Kernel Transformer Networks for Compact Spherical Convolution
Yu-Chuan Su, Kristen Grauman
Visual Question: Predicting If a Crowd Will Agree on the Answer
Danna Gurari, Kristen Grauman
Human Action Anticipation: A Survey
Bolin Lai, Sam Toyer, Tushar Nagarajan +7
Vid2Coach: Transforming How-To Videos into Task Assistants
Mina Huh, Zihui Xue, Ujjaini Das +3
Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual Navigation
Ziad Al-Halah, Santhosh K. Ramakrishnan, Kristen Grauman
Fashion++: Minimal Edits for Outfit Improvement
Wei-Lin Hsiao, Isay Katsman, Chao-Yuan Wu +2
Learning Object State Changes in Videos: An Open-World Perspective
Zihui Xue, Kumar Ashutosh, Kristen Grauman
Learning to Set Waypoints for Audio-Visual Navigation
Changan Chen, Sagnik Majumder, Ziad Al-Halah +3
From Paris to Berlin: Discovering Fashion Style Influences Around the World
Ziad Al-Halah, Kristen Grauman
Pixel Objectness: Learning to Segment Generic Objects Automatically in Images and Videos
Bo Xiong, Suyog Dutt Jain, Kristen Grauman
Multiview Pseudo-Labeling for Semi-supervised Learning from Video
Bo Xiong, Haoqi Fan, Kristen Grauman +1
Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment
Zihui Xue, Kristen Grauman
Incentivizing Vision Language Models to Search for Long Video Question Answering
Harsh Goel, S P Sharan, Sahil Shah +4
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
Sagnik Majumder, Ziad Al-Halah, Kristen Grauman
Detours for Navigating Instructional Videos
Kumar Ashutosh, Zihui Xue, Tushar Nagarajan +1
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
Joungbin An, Kristen Grauman
Fashion Forward: Forecasting Visual Style in Fashion
Ziad Al-Halah, Rainer Stiefelhagen, Kristen Grauman
Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction
Changan Chen, Jordi Ramos, Anshul Tomar +1
Learning Spherical Convolution for Fast Features from 360° Imagery
Yu-Chuan Su, Kristen Grauman
Learning to Separate Object Sounds by Watching Unlabeled Video
Ruohan Gao, Rogerio Feris, Kristen Grauman