7 papers
Xray-Visual Models: Scaling Vision models on Industry Scale Data
Shlok Mishra, Tsung-Yu Lin, Linda Wang +24
We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 b…
Principled Synthetic Data Enables the First Scaling Laws for LLMs in Recommendation
Benyu Zhang, Qiang Zhang, Jianpeng Cheng +10
Large Language Models (LLMs) represent a promising frontier for recommender systems, yet their development has been impeded by the absence of predictable scaling laws, which are cr…
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
Chao Li, Tianhong Li, Sai Vidyaranya Nuthalapati +9
Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives demand opposing masking regimes:…
Think Then Embed: Generative Context Improves Multimodal Embedding
Xuanming Cui, Jianpeng Cheng, Hong-you Chen +11
There is a growing interest in Universal Multimodal Embeddings (UME), where models are required to generate task-specific representations. While recent studies show that Multimodal…
Reason to Contrast: A Cascaded Multimodal Retrieval Framework
Xuanming Cui, Hong-You Chen, Hao Yu +10
Traditional multimodal retrieval systems rely primarily on bi-encoder architectures, where performance is closely tied to embedding dimensionality. Recent work, Think-Then-Embed (T…
A Systematic Characterization of LLM Inference on GPUs
Haonan Wang, Xuxin Xiao, Mingyu Yan +8
This work presents a systematic characterization of Large Language Model (LLM) inference to address fragmented understanding. Through comprehensive experiments, we establish a four…