6 papers
DisasterInsight: A Multimodal Benchmark for Function-Aware and Grounded Disaster Assessment
Sara Tehrani, Yonghao Xu, Leif Haglund +2
Timely interpretation of satellite imagery is critical for disaster response, yet existing vision-language benchmarks for remote sensing largely focus on coarse labels and image-le…
VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
Emanuel Sánchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan +2
Satellite imagery differs fundamentally from natural images: its aerial viewpoint, very high resolution, diverse scale variations, and abundance of small objects demand both region…
Out-of-Distribution Segmentation via Wasserstein-Based Evidential Uncertainty
Arnold Brosch, Abdelrahman Eldesokey, Michael Felsberg +1
Deep neural networks achieve superior performance in semantic segmentation, but are limited to a predefined set of classes, which leads to failures when they encounter unknown obje…
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz +26
Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understand and generate descriptions of visual content. Most existing LMMs are in Englis…
From Missing Pieces to Masterpieces: Image Completion with Context-Adaptive Diffusion
Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou +3
Image completion is a challenging task, particularly when ensuring that generated content seamlessly integrates with existing parts of an image. While recent diffusion models have…
All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages
Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana +66
Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cul…