18 citations · 44 across the 4 of their papers we have counts for
6 papers · 1 filter
HourVideo: 1-Hour Video-Language Understanding
Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic +7
We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, track…
Holistic Evaluation of Text-To-Image Models
Tony Lee, Michihiro Yasunaga, Chenlin Meng +15
The stunning qualitative improvement of recent text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding…
Siamese Masked Autoencoders
Agrim Gupta, Jiajun Wu, Jia Deng +1
Establishing correspondence between images or scenes is a significant challenge in computer vision, especially given occlusions, viewpoint changes, and varying object appearances.…
Image Generation from Scene Graphs
Justin Johnson, Agrim Gupta, Li Fei-Fei
To truly understand the visual world our models should be able not only to recognize images but also generate them. To this end, there has been exciting recent progress on generati…
Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks
Agrim Gupta, Justin Johnson, Li Fei-Fei +2
Understanding human motion behavior is critical for autonomous moving platforms (like self-driving cars and social robots) if they are to navigate human-centric environments. This…
Characterizing and Improving Stability in Neural Style Transfer
Agrim Gupta, Justin Johnson, Alexandre Alahi +1
Recent progress in style transfer on images has focused on improving the quality of stylized images and speed of methods. However, real-time methods are highly unstable resulting i…