3 papers
cs.CV2025
Inference Compute-Optimal Video Vision Language Models
Peiqi Wang, ShengYun Peng, Xuewen Zhang +5
This work investigates the optimal allocation of inference compute across three key scaling factors in video vision language models: language model size, frame count, and the numbe…
cs.IR2025
Towards An Efficient LLM Training Paradigm for CTR Prediction
Allen Lin, Renqin Cai, Yun He +5
Large Language Models (LLMs) have demonstrated tremendous potential as the next-generation ranking-based recommendation system. Many recent works have shown that LLMs can significa…
cs.CV2024
CompCap: Improving Multimodal Large Language Models with Composite Captions
Xiaohui Chen, Satya Narayan Shukla, Mahmoud Azab +8
How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as…