6 papers
Xray-Visual Models: Scaling Vision models on Industry Scale Data
Shlok Mishra, Tsung-Yu Lin, Linda Wang +24
We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 b…
Leveraging Data to Say No: Memory Augmented Plug-and-Play Selective Prediction
Aditya Sarkar, Yi Li, Jiacheng Cheng +2
Selective prediction aims to endow predictors with a reject option, to avoid low confidence predictions. However, existing literature has primarily focused on closed-set tasks, suc…
Think Then Embed: Generative Context Improves Multimodal Embedding
Xuanming Cui, Jianpeng Cheng, Hong-you Chen +11
There is a growing interest in Universal Multimodal Embeddings (UME), where models are required to generate task-specific representations. While recent studies show that Multimodal…
Socratic Students: Teaching Language Models to Learn by Asking Questions
Rajeev Bhatt Ambati, Tianyi Niu, Aashu Singh +3
Large language Models (LLMs) are usually used to answer questions, but many high-stakes applications (e.g., tutoring, clinical support) require the complementary skill of asking qu…
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
Yanlai Yang, Zhuokai Zhao, Satya Narayan Shukla +4
Multimodal large language models (MLLMs) have made significant progress in visual-language reasoning, but their ability to efficiently handle long videos remains limited. Despite r…
Transfer between Modalities with MetaQueries
Xichen Pan, Satya Narayan Shukla, Aashu Singh +9
Unified multimodal models aim to integrate understanding (text output) and generation (pixel output), but aligning these different modalities within a single architecture often dem…