4 papers
CoFi-UCGen: Coarse-to-Fine Unsupervised Conditional Generation without Label Priors
Shengxi Li, Zhaokun Hu, Ce Zheng +3
Unsupervised conditional image generation (UCGen) aims to control generation without relying on manually annotated labels, yet remains challenging due to unstructured semantic repr…
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX
Liuyue Xie, Avik Kuthiala, George Z. Wei +12
We introduce MAVERIX (Multimodal audiovisual Evaluation and Recognition IndeX), a unified benchmark to probe the video understanding in multimodal LLMs, encompassing video, audio,…
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…
VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models
Harshit, Tolga Tasdizen
The recent developments in deep learning led to the integration of natural language processing (NLP) with computer vision, resulting in powerful integrated Vision and Language Mode…