From the 2 of 14 linked papers with an AI index.
2 citations · 2 across the 8 of their papers we have counts for
7 papers · 1 filter
Xray-Visual Models: Scaling Vision models on Industry Scale Data
Shlok Mishra, Tsung-Yu Lin, Linda Wang +24
We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 b…
RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space
Xichen Pan, Aashu Singh, Satya Narayan Shukla +3
Large language models (LLMs) are widely used in text-to-image (T2I) systems, but they are typically limited to text encoding, while denoising is handled by newly trained generative…
An Attribute-Based Measure of Video Complexity
Aditya Sarkar, Yi Li, Zihao Wang +6
A new framework for the estimation of the complexity posed by video-question pairs to video-LLMs, Video Attribute-Based Complexity (VideoABC), is proposed. Video complexity is defi…
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
Chao Li, Tianhong Li, Sai Vidyaranya Nuthalapati +9
Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives demand opposing masking regimes:…
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
Yanlai Yang, Zhuokai Zhao, Satya Narayan Shukla +4
Multimodal large language models (MLLMs) have made significant progress in visual-language reasoning, but their ability to efficiently handle long videos remains limited. Despite r…
Transfer between Modalities with MetaQueries
Xichen Pan, Satya Narayan Shukla, Aashu Singh +9
Unified multimodal models aim to integrate understanding (text output) and generation (pixel output), but aligning these different modalities within a single architecture often dem…