activity
20242026
collaborators

6 papers

cs.CV2026

DisasterInsight: A Multimodal Benchmark for Function-Aware and Grounded Disaster Assessment

Sara Tehrani, Yonghao Xu, Leif Haglund +2

Timely interpretation of satellite imagery is critical for disaster response, yet existing vision-language benchmarks for remote sensing largely focus on coarse labels and image-le…

cs.CV2025

VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing

Emanuel Sánchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan +2

Satellite imagery differs fundamentally from natural images: its aerial viewpoint, very high resolution, diverse scale variations, and abundance of small objects demand both region…

cs.CV2025

Out-of-Distribution Segmentation via Wasserstein-Based Evidential Uncertainty

Arnold Brosch, Abdelrahman Eldesokey, Michael Felsberg +1

Deep neural networks achieve superior performance in semantic segmentation, but are limited to a predefined set of classes, which leads to failures when they encounter unknown obje…

cs.CL2025

A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz +26

Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understand and generate descriptions of visual content. Most existing LMMs are in Englis…

cs.CV2025

From Missing Pieces to Masterpieces: Image Completion with Context-Adaptive Diffusion

Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou +3

Image completion is a challenging task, particularly when ensuring that generated content seamlessly integrates with existing parts of an image. While recent diffusion models have…

cs.CV2024

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana +66

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cul…