8 papers · 1 filter
Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation
Kun Li, Shengxi Gui, Francesco Nex +1
Referring Remote Sensing Image Segmentation (RRSIS) seeks to localize and segment the target object or region specified by a natural language expression in a remote sensing image.…
Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question Answering
Kun Li, Michael Ying Yang, Sami Sebastian Brandt
Audio--Visual Question Answering (AVQA) is a challenging multimodal task that requires jointly reasoning over audio, visual, and textual information in a given video to answer natu…
DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion
Mengmeng Liu, Michael Ying Yang, Jiuming Liu +5
Visual-LiDAR odometry is a critical component for autonomous system localization, yet achieving high accuracy and strong robustness remains a challenge. Traditional approaches comm…
Multimodal Rationales for Explainable Visual Question Answering
Kun Li, George Vosselman, Michael Ying Yang
Visual Question Answering (VQA) is a challenging task of predicting the answer to a question about the content of an image. Prior works directly evaluate the answering models by si…
Scale-wise Bidirectional Alignment Network for Referring Remote Sensing Image Segmentation
Kun Li, George Vosselman, Michael Ying Yang
The goal of referring remote sensing image segmentation (RRSIS) is to extract specific pixel-level regions within an aerial image via a natural language expression. Recent advancem…
Planner3D: LLM-enhanced graph prior meets 3D indoor scene explicit regularization
Yao Wei, Martin Renqiang Min, George Vosselman +2
Compositional 3D scene synthesis has diverse applications across a spectrum of industries such as robotics, films, and video games, as it closely mirrors the complexity of real-wor…