5 papers
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
Ailar Mahdizadeh, Puria Azadi, Muchen Li +2
Streaming video understanding with large vision-language models (VLMs) requires a compact memory that can support future reasoning over an ever-growing visual history. A common sol…
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
Tanzila Rahman, Renjie Liao, Leonid Sigal
Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, an…
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
Jia Jun Cheng Xian, Muchen Li, Haotian Yang +4
Recent advances in diffusion-based text-to-image (T2I) models have led to remarkable success in generating high-quality images from textual prompts. However, ensuring accurate alig…
Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
Bicheng Xu, Qi Yan, Renjie Liao +2
We introduce a framework for joint grounded scene graph - image generation, a challenging task involving high-dimensional, multi-modal structured data. To effectively model this co…
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
Sadegh Mahdavi, Muchen Li, Kaiwen Liu +3
Advances in Large Language Models (LLMs) have sparked interest in their ability to solve Olympiad-level math problems. However, the training and evaluation of these models are cons…