6 papers
Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation
Kun Li, Shengxi Gui, Francesco Nex +1
Referring Remote Sensing Image Segmentation (RRSIS) seeks to localize and segment the target object or region specified by a natural language expression in a remote sensing image.…
AccioScene: Compositional 3D Scene Generation via Graph Diffusion and Interaction-driven Critics
Yao Wei, Matteo Toso, Pietro Morerio +3
This paper presents a framework for generating 3D indoor scenes from text prompts. Existing methods often formulate scene synthesis as an object layout prediction problem condition…
Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question Answering
Kun Li, Michael Ying Yang, Sami Sebastian Brandt
Audio--Visual Question Answering (AVQA) is a challenging multimodal task that requires jointly reasoning over audio, visual, and textual information in a given video to answer natu…
DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion
Mengmeng Liu, Michael Ying Yang, Jiuming Liu +5
Visual-LiDAR odometry is a critical component for autonomous system localization, yet achieving high accuracy and strong robustness remains a challenge. Traditional approaches comm…
Multimodal Rationales for Explainable Visual Question Answering
Kun Li, George Vosselman, Michael Ying Yang
Visual Question Answering (VQA) is a challenging task of predicting the answer to a question about the content of an image. Prior works directly evaluate the answering models by si…
Scale-wise Bidirectional Alignment Network for Referring Remote Sensing Image Segmentation
Kun Li, George Vosselman, Michael Ying Yang
The goal of referring remote sensing image segmentation (RRSIS) is to extract specific pixel-level regions within an aerial image via a natural language expression. Recent advancem…