NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (153)

cs.CV2023

Egocentric Video Task Translation @ Ego4D Challenge 2022

Zihui Xue, Yale Song, Kristen Grauman +1

cs.CV2022

PONI: Potential Functions for ObjectGoal Navigation with Interaction-free Learning

Santhosh Kumar Ramakrishnan, Devendra Singh Chaplot, Ziad Al-Halah +2

cs.CV2016

Learning image representations tied to ego-motion

Dinesh Jayaraman, Kristen Grauman

cs.CV2020

Learning Affordance Landscapes for Interaction Exploration in 3D Environments

Tushar Nagarajan, Kristen Grauman

cs.CV2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

Kristen Grauman, Andrew Westbury, Eugene Byrne +82

cs.CV2022

Visual Acoustic Matching

Changan Chen, Ruohan Gao, Paul Calamia +1

cs.CV2022

Retrospectives on the Embodied AI Workshop

Matt Deitke, Dhruv Batra, Yonatan Bisk +36

cs.CV2024

SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos

Changan Chen, Kumar Ashutosh, Rohit Girdhar +2

cs.CV2023

EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding

Shuhan Tan, Tushar Nagarajan, Kristen Grauman

cs.CV2018

ShapeCodes: Self-Supervised Feature Learning by Lifting Views to Viewgrids

Dinesh Jayaraman, Ruohan Gao, Kristen Grauman

cs.CV2019

Co-Separating Sounds of Visual Objects

Ruohan Gao, Kristen Grauman

cs.CV2026

Seeing without Pixels: Perception from Camera Trajectories

Zihui Xue, Kristen Grauman, Dima Damen +2

cs.CV2023

Novel-View Acoustic Synthesis

Changan Chen, Alexander Richard, Roman Shapovalov +4

cs.CV2023

Single-Stage Visual Query Localization in Egocentric Videos

Hanwen Jiang, Santhosh Kumar Ramakrishnan, Kristen Grauman

cs.CV2025

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

Mi Luo, Zihui Xue, Alex Dimakis +1

cs.CV2019

Grounded Human-Object Interaction Hotspots from Video (Extended Abstract)

Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman

cs.CV2025

Learning Skill-Attributes for Transferable Assessment in Video

Kumar Ashutosh, Kristen Grauman

cs.CV2017

Semantic Jitter: Dense Supervision for Visual Comparisons via Synthetic Images

Aron Yu, Kristen Grauman

cs.CV2022

Active Audio-Visual Separation of Dynamic Sound Sources

Sagnik Majumder, Kristen Grauman

cs.CV2020

SoundSpaces: Audio-Visual Navigation in 3D Environments

Changan Chen, Unnat Jain, Carl Schissler +5

cs.CV2016

Look-ahead before you leap: end-to-end active recognition by forecasting the effect of motion

Dinesh Jayaraman, Kristen Grauman

cs.CV2025

Seeing the Arrow of Time in Large Multimodal Models

Zihui Xue, Mi Luo, Kristen Grauman

cs.CV2020

Audio-Visual Floorplan Reconstruction

Senthil Purushwalkam, Sebastian Vicenc Amengual Gari, Vamsi Krishna Ithapu +4

cs.CV2018

Sidekick Policy Learning for Active Visual Exploration

Santhosh K. Ramakrishnan, Kristen Grauman

stat.ML2014

Large-Margin Determinantal Point Processes

Boqing Gong, Wei-lun Chao, Kristen Grauman +1

cs.CV2023

HierVL: Learning Hierarchical Video-Language Embeddings

Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani +1

cs.CV2026

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models

Ami Baid, Zihui Xue, Kristen Grauman

cs.CV2024

HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness

Zihui Xue, Mi Luo, Changan Chen +1

cs.CV2018

Compare and Contrast: Learning Prominent Visual Differences

Steven Chen, Kristen Grauman

cs.CV2023

Chat2Map: Efficient Scene Mapping from Multi-Ego Conversations

Sagnik Majumder, Hao Jiang, Pierre Moulon +4

cs.CV2017

Predicting Foreground Object Ambiguity and Efficiently Crowdsourcing the Segmentation(s)

Danna Gurari, Kun He, Bo Xiong +6

cs.CV2015

WhittleSearch: Interactive Image Search with Relative Attribute Feedback

Adriana Kovashka, Devi Parikh, Kristen Grauman

cs.CV2026

Human detectors are surprisingly powerful reward models

Kumar Ashutosh, XuDong Wang, Xi Yin +4

cs.CV2020

ViBE: Dressing for Diverse Body Shapes

Wei-Lin Hsiao, Kristen Grauman

cs.CV2019

BlockDrop: Dynamic Inference Paths in Residual Networks

Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar +4

cs.CV2025

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi +26

cs.CV2020

Listen to Look: Action Recognition by Previewing Audio

Ruohan Gao, Tae-Hyun Oh, Kristen Grauman +1

cs.CV2016

Object-Centric Representation Learning from Unlabeled Videos

Ruohan Gao, Dinesh Jayaraman, Kristen Grauman

cs.CV2016

Pano2Vid: Automatic Cinematography for Watching 360$^{\circ}$ Videos

Yu-Chuan Su, Dinesh Jayaraman, Kristen Grauman

cs.CV2023

What You Say Is What You Show: Visual Narration Detection in Instructional Videos

Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani +1

cs.CV2024

Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos

Mi Luo, Zihui Xue, Alex Dimakis +1

cs.CV2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98

econ.EM2024

Shapes as Product Differentiation: Neural Network Embedding in the Analysis of Markets for Fonts

Sukjin Han, Eric H. Schulman, Kristen Grauman +1

cs.CV2026

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding

Joungbin An, Agrim Jain, Kristen Grauman

cs.CV2020

VisualEchoes: Spatial Image Representation Learning through Echolocation

Ruohan Gao, Changan Chen, Ziad Al-Halah +2

cs.CV2016

Efficient Activity Detection in Untrimmed Video with Max-Subgraph Search

Chao-Yeh Chen, Kristen Grauman

cs.CV2018

Snap Angle Prediction for 360$^{\circ}$ Panoramas

Bo Xiong, Kristen Grauman

cs.CV2021

From Culture to Clothing: Discovering the World Events Behind A Century of Fashion Images

Wei-Lin Hsiao, Kristen Grauman

cs.CV2026

Personal Visual Context Learning in Large Multimodal Models

Zihui Xue, Ami Baid, Sangho Kim +2

cs.CV2025

Progress-Aware Video Frame Captioning

Zihui Xue, Joungbin An, Xitong Yang +1

cs.CY2018

Video Analysis for Body-worn Cameras in Law Enforcement

Jason J. Corso, Alexandre Alahi, Kristen Grauman +4

cs.CV2017

Learning Compressible 360° Video Isomers

Yu-Chuan Su, Kristen Grauman

cs.CV2018

Im2Flow: Motion Hallucination from Static Images for Action Recognition

Ruohan Gao, Bo Xiong, Kristen Grauman

cs.CV2023

NaQ: Leveraging Narrations as Queries to Supervise Episodic Memory

Santhosh Kumar Ramakrishnan, Ziad Al-Halah, Kristen Grauman

cs.CV2020

Discovering Underground Maps from Fashion

Utkarsh Mall, Kavita Bala, Tamara Berg +1

cs.CV2021

Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos

Yanghao Li, Tushar Nagarajan, Bo Xiong +1

cs.CV2017

FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos

Suyog Dutt Jain, Bo Xiong, Kristen Grauman

cs.RO2022

DexVIP: Learning Dexterous Grasping with Human Hand Pose Priors from Video

Priyanka Mandikal, Kristen Grauman

cs.CV2020

Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback

Hui Wu, Yupeng Gao, Xiaoxiao Guo +4

cs.CV2018

SpotTune: Transfer Learning through Adaptive Fine-tuning

Yunhui Guo, Honghui Shi, Abhishek Kumar +3

cs.CV2019

Emergence of Exploratory Look-Around Behaviors through Active Observation Completion

Santhosh K. Ramakrishnan, Dinesh Jayaraman, Kristen Grauman

cs.CV2017

Learning the Latent "Look": Unsupervised Discovery of a Style-Coherent Embedding from Fashion Images

Wei-Lin Hsiao, Kristen Grauman

cs.SD2023

Learning Audio-Visual Dereverberation

Changan Chen, Wei Sun, David Harwath +1

cs.CV2020

An Exploration of Embodied Visual Exploration

Santhosh K. Ramakrishnan, Dinesh Jayaraman, Kristen Grauman

cs.CV2020

You2Me: Inferring Body Pose in Egocentric Video via First and Second Person Interactions

Evonne Ng, Donglai Xiang, Hanbyul Joo +1

cs.CV2015

Predicting Important Objects for Egocentric Video Summarization

Yong Jae Lee, Kristen Grauman

cs.CV2020

Occupancy Anticipation for Efficient Exploration and Navigation

Santhosh K. Ramakrishnan, Ziad Al-Halah, Kristen Grauman

cs.CV2016

Detangling People: Individuating Multiple Close People and Their Body Parts via Region Assembly

Hao Jiang, Kristen Grauman

cs.CV2016

Leaving Some Stones Unturned: Dynamic Feature Prioritization for Activity Detection in Streaming Video

Yu-Chuan Su, Kristen Grauman

cs.CV2020

EGO-TOPO: Environment Affordances from Egocentric Video

Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer +1

cs.CV2018

VizWiz Grand Challenge: Answering Visual Questions from Blind People

Danna Gurari, Qing Li, Abigale J. Stangl +5

cs.CV2025

Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos

Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah +1

cs.CV2019

Grounded Human-Object Interaction Hotspots from Video

Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman

cs.CV2019

Extreme Relative Pose Estimation for RGB-D Scans via Scene Completion

Zhenpei Yang, Jeffrey Z. Pan, Linjie Luo +3

cs.CV2025

ExpertAF: Expert Actionable Feedback from Video

Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos +2

cs.CV2022

Egocentric Activity Recognition and Localization on a 3D Map

Miao Liu, Lingni Ma, Kiran Somasundaram +4

cs.SD2022

Few-Shot Audio-Visual Learning of Environment Acoustics

Sagnik Majumder, Changan Chen, Ziad Al-Halah +1

cs.LG2023

A Domain-Agnostic Approach for Characterization of Lifelong Learning Systems

Megan M. Baker, Alexander New, Mario Aguilar-Simon +44

cs.CV2020

Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias

Krishna Kumar Singh, Dhruv Mahajan, Kristen Grauman +3

cs.RO2026

Mash, Spread, Slice! Learning to Manipulate Object States via Visual Spatial Progress

Priyanka Mandikal, Jiaheng Hu, Shivin Dass +3

cs.CV2019

Kernel Transformer Networks for Compact Spherical Convolution

Yu-Chuan Su, Kristen Grauman

cs.AI2016

Visual Question: Predicting If a Crowd Will Agree on the Answer

Danna Gurari, Kristen Grauman

cs.CV2024

Human Action Anticipation: A Survey

Bolin Lai, Sam Toyer, Tushar Nagarajan +7

cs.HC2025

Vid2Coach: Transforming How-To Videos into Task Assistants

Mina Huh, Zihui Xue, Ujjaini Das +3

cs.CV2022

Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual Navigation

Ziad Al-Halah, Santhosh K. Ramakrishnan, Kristen Grauman

cs.CV2019

Fashion++: Minimal Edits for Outfit Improvement

Wei-Lin Hsiao, Isay Katsman, Chao-Yuan Wu +2

cs.CV2024

Learning Object State Changes in Videos: An Open-World Perspective

Zihui Xue, Kumar Ashutosh, Kristen Grauman

cs.CV2021

Learning to Set Waypoints for Audio-Visual Navigation

Changan Chen, Sagnik Majumder, Ziad Al-Halah +3

cs.CV2020

From Paris to Berlin: Discovering Fashion Style Influences Around the World

Ziad Al-Halah, Kristen Grauman

cs.CV2018

Pixel Objectness: Learning to Segment Generic Objects Automatically in Images and Videos

Bo Xiong, Suyog Dutt Jain, Kristen Grauman

cs.CV2021

Multiview Pseudo-Labeling for Semi-supervised Learning from Video

Bo Xiong, Haoqi Fan, Kristen Grauman +1

cs.CV2023

Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment

Zihui Xue, Kristen Grauman

cs.CV2026

Incentivizing Vision Language Models to Search for Long Video Question Answering

Harsh Goel, S P Sharan, Sahil Shah +4

cs.CV2024

Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos

Sagnik Majumder, Ziad Al-Halah, Kristen Grauman

cs.CV2024

Detours for Navigating Instructional Videos

Kumar Ashutosh, Zihui Xue, Tushar Nagarajan +1

cs.CV2026

HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling

Joungbin An, Kristen Grauman

cs.CV2020

Fashion Forward: Forecasting Visual Style in Fashion

Ziad Al-Halah, Rainer Stiefelhagen, Kristen Grauman

cs.SD2024

Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction

Changan Chen, Jordi Ramos, Anshul Tomar +1

cs.CV2018

Learning Spherical Convolution for Fast Features from 360° Imagery

Yu-Chuan Su, Kristen Grauman

cs.CV2018

Learning to Separate Object Sounds by Watching Unlabeled Video

Ruohan Gao, Rogerio Feris, Kristen Grauman