118 citations · 475 across the 26 of their papers we have counts for
11 papers · 1 filter
Grounded Human-Object Interaction Hotspots from Video
Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman
Learning how to interact with objects is an important step towards embodied visual intelligence, but existing techniques suffer from heavy supervision or sensing requirements. We p…
2.5D Visual Sound
Ruohan Gao, Kristen Grauman
Binaural audio provides a listener with 3D sound sensation, allowing a rich perceptual experience of the scene. However, binaural recordings are scarcely available and require nont…
Kernel Transformer Networks for Compact Spherical Convolution
Yu-Chuan Su, Kristen Grauman
Ideally, 360° imagery could inherit the deep convolutional neural networks (CNNs) already trained with great success on perspective projection images. However, existing methods to…
SpotTune: Transfer Learning through Adaptive Fine-tuning
Yunhui Guo, Honghui Shi, Abhishek Kumar +3
Transfer learning, which allows a source task to affect the inductive bias of the target task, is widely used in computer vision. The typical way of conducting transfer learning wi…
Pixel Objectness: Learning to Segment Generic Objects Automatically in Images and Videos
Bo Xiong, Suyog Dutt Jain, Kristen Grauman
We propose an end-to-end learning framework for segmenting generic objects in both images and videos. Given a novel image or video, our approach produces a pixel-level mask for all…
Sidekick Policy Learning for Active Visual Exploration
Santhosh K. Ramakrishnan, Kristen Grauman
We consider an active visual exploration scenario, where an agent must intelligently select its camera motions to efficiently reconstruct the full environment from only a limited s…