47 citations · 47 across the 1 of their papers we have counts for
8 papers
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
VideoPrism: A Foundational Visual Encoder for Video Understanding
Long Zhao, Nitesh B. Gundavarapu, Liangzhe Yuan +16
We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus…
Epsilon-VAE: Denoising as Visual Decoding
Long Zhao, Sanghyun Woo, Ziyu Wan +6
In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data,…
SANPO: A Scene Understanding, Accessibility and Human Navigation Dataset
Sagar M. Waghmare, Kimberly Wilber, Dave Hawkey +9
Vision is essential for human navigation. The World Health Organization (WHO) estimates that 43.3 million people were blind in 2020, and this number is projected to reach 61 millio…
Video Creation by Demonstration
Yihong Sun, Hao Zhou, Liangzhe Yuan +7
We explore a novel video creation experience, namely Video Creation by Demonstration. Given a demonstration video and a context image from a different scene, we generate a physical…
VideoGLUE: Video General Understanding Evaluation of Foundation Models
Liangzhe Yuan, Nitesh Bharadwaj Gundavarapu, Long Zhao +14
We evaluate the video understanding capabilities of existing foundation models (FMs) using a carefully designed experiment protocol consisting of three hallmark tasks (action recog…