18 citations · 44 across the 6 of their papers we have counts for
8 papers · 1 filter
Movie Gen: A Cast of Media Foundation Models
Adam Polyak, Amit Zohar, Andrew Brown +85
We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabili…
DISGO: Automatic End-to-End Evaluation for Scene Text OCR
Mei-Yuh Hwang, Yangyang Shi, Ankit Ramchandani +6
This paper discusses the challenges of optical character recognition (OCR) on natural scenes, which is harder than OCR on documents due to the wild content and various image backgr…
Episodic Memory Question Answering
Samyak Datta, Sameer Dharur, Vincent Cartillier +4
Egocentric augmented reality devices such as wearable glasses passively capture visual data as a human wearer tours a home environment. We envision a scenario wherein the human com…
Integrating Egocentric Localization for More Realistic Point-Goal Navigation Agents
Samyak Datta, Oleksandr Maksymets, Judy Hoffman +3
Recent work has presented embodied agents that can navigate to point-goal targets in novel indoor environments with near-perfect accuracy. However, these agents are equipped with i…
Embodied Question Answering in Photorealistic Environments with Point Cloud Perception
Erik Wijmans, Samyak Datta, Oleksandr Maksymets +6
To help bridge the gap between internet vision-style problems and the goal of vision for embodied perception we instantiate a large-scale navigation task -- Embodied Question Answe…
Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption Alignment
Samyak Datta, Karan Sikka, Anirban Roy +3
We address the problem of grounding free-form textual phrases by using weak supervision from image-caption pairs. We propose a novel end-to-end model that uses caption-to-image ret…