20 papers
Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification
Maryam Gholami Shiri, Eva Tuba, Sašo Džeroski +2
Benchmarking deep learning (DL) models for multi-label classification (MLC) of remote sensing images (RSI) typically yields rankings that do not generalize beyond the evaluated dat…
Unsupervised Multi-kernel Learning for Automated Algorithm Selection
Yihang Lu, Tome Eftimov, Carola Doerr
Automated algorithm selection in black-box optimization typically relies on supervised models that map landscape features to algorithm performance labels. Such models are costly to…
Graph Instance Landscapes: When Structural Similarity Does (Not) Reflect Shortest-Path Performance
Maryam Gholami Shiri, Ivana Krminac, Marko DjukanoviÄ +3
Benchmarking shortest-path algorithms is commonly based on aggregate performance over heterogeneous graph sets, which limits insight into how different search paradigms react to in…
Evaluating Real-World Generalizability of Algorithm Selection Models
Gjorgjina Cenikj, Jakub Kudela, Eva Tuba +1
Algorithm Selection (AS) aims to automatically identify the most suitable optimization algorithm for a given problem instance by leveraging measurable problem characteristics and h…
Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs
Denica Kjorvezir, Marko DjukanoviÄ, Ana Gjorgjevikj +2
Evaluating large language models (LLMs) across comprehensive benchmarks is expensive and time-consuming. We propose a graph-based prompt selection framework that models each benchm…
On the Robustness of Multilingual Text Embedding Rankings Across Learning Tasks, Languages, and Benchmark Datasets
Ana Gjorgjevikj, Barbara KorouÅ¡iÄ Seljak, Tome Eftimov
Large-scale multilingual text embedding models play crucial role in both research and industry, yet their behavior in language-specific, multi-task settings remains insufficiently…