4 papers · 1 filter
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
Ailar Mahdizadeh, Puria Azadi, Muchen Li +2
Streaming video understanding with large vision-language models (VLMs) requires a compact memory that can support future reasoning over an ever-growing visual history. A common sol…
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
Tanzila Rahman, Renjie Liao, Leonid Sigal
Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, an…
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
Jia Jun Cheng Xian, Muchen Li, Haotian Yang +4
Recent advances in diffusion-based text-to-image (T2I) models have led to remarkable success in generating high-quality images from textual prompts. However, ensuring accurate alig…
Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
Bicheng Xu, Qi Yan, Renjie Liao +2
We introduce a framework for joint grounded scene graph - image generation, a challenging task involving high-dimensional, multi-modal structured data. To effectively model this co…