6 papers
Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question Answering
Kun Li, Michael Ying Yang, Sami Sebastian Brandt
Audio--Visual Question Answering (AVQA) is a challenging multimodal task that requires jointly reasoning over audio, visual, and textual information in a given video to answer natu…
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
Mingjie Xu, Jinpeng Chen, Yuzhi Zhao +12
Multimodal large language models (MLLMs) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding.…
MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
Qian Liang, Yujia Wu, Kuncheng Li +4
Multimodal Large Language Models (MLLMs) with unified architectures excel across a wide range of vision-language tasks, yet aligning them with personalized image generation remains…
From Sound to Sight: Towards AI-authored Music Videos
Leo Vitasovic, Stella GraÃhof, Agnes Mercedes Kloft +6
Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for…
Multimodal Rationales for Explainable Visual Question Answering
Kun Li, George Vosselman, Michael Ying Yang
Visual Question Answering (VQA) is a challenging task of predicting the answer to a question about the content of an image. Prior works directly evaluate the answering models by si…
Scale-wise Bidirectional Alignment Network for Referring Remote Sensing Image Segmentation
Kun Li, George Vosselman, Michael Ying Yang
The goal of referring remote sensing image segmentation (RRSIS) is to extract specific pixel-level regions within an aerial image via a natural language expression. Recent advancem…