collaborators

20 papers

cs.LG2026

Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification

Maryam Gholami Shiri, Eva Tuba, Sašo Džeroski +2

Benchmarking deep learning (DL) models for multi-label classification (MLC) of remote sensing images (RSI) typically yields rankings that do not generalize beyond the evaluated dat…

cs.LG2026

Unsupervised Multi-kernel Learning for Automated Algorithm Selection

Yihang Lu, Tome Eftimov, Carola Doerr

Automated algorithm selection in black-box optimization typically relies on supervised models that map landscape features to algorithm performance labels. Such models are costly to…

cs.SI2026

Graph Instance Landscapes: When Structural Similarity Does (Not) Reflect Shortest-Path Performance

Maryam Gholami Shiri, Ivana Krminac, Marko Djukanović +3

Benchmarking shortest-path algorithms is commonly based on aggregate performance over heterogeneous graph sets, which limits insight into how different search paradigms react to in…

cs.LG2026

Evaluating Real-World Generalizability of Algorithm Selection Models

Gjorgjina Cenikj, Jakub Kudela, Eva Tuba +1

Algorithm Selection (AS) aims to automatically identify the most suitable optimization algorithm for a given problem instance by leveraging measurable problem characteristics and h…

cs.CL2026

Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs

Denica Kjorvezir, Marko Djukanović, Ana Gjorgjevikj +2

Evaluating large language models (LLMs) across comprehensive benchmarks is expensive and time-consuming. We propose a graph-based prompt selection framework that models each benchm…

cs.CL2026

On the Robustness of Multilingual Text Embedding Rankings Across Learning Tasks, Languages, and Benchmark Datasets

Ana Gjorgjevikj, Barbara Koroušić Seljak, Tome Eftimov

Large-scale multilingual text embedding models play crucial role in both research and industry, yet their behavior in language-specific, multi-task settings remains insufficiently…