1 citations · 1 across the 2 of their papers we have counts for
4 papers
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
Madhuri Shanbhogue, Zhe Li, Shanfeng Zhang +86
We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified representation space. We leverage t…
Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion
Chien Van Nguyen, Chaitra Hegde, Van Cuong Pham +3
We introduce Orthrus, a simple and efficient dual-architecture framework that unifies the exact generation fidelity of autoregressive Large Language Models (LLMs) with the high-spe…
Extending Video Masked Autoencoders to 128 frames
Nitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal +8
Video understanding has witnessed significant progress with recent video foundation models demonstrating strong performance owing to self-supervised pre-training objectives; Masked…
Feasibility of assessing cognitive impairment via distributed camera network and privacy-preserving edge computing
Chaitra Hegde, Yashar Kiarashi, Allan I Levey +3
INTRODUCTION: Mild cognitive impairment (MCI) is characterized by a decline in cognitive functions beyond typical age and education-related expectations. Since, MCI has been linked…