10 papers
Agreement in Representation Space for Open-Ended Self-Consistency
Paula Ontalvilla, Gorka Azkune, Aitor Ormazabal
Self-consistency improves LLM reasoning by sampling multiple outputs and selecting the most consistent answer, but existing formulations largely rely on exact matching and therefor…
Mixture of Predefined Experts: Maximizing Data Usage on Vertical Federated Learning
Jon Irureta, Gorka Azkune, Jon Imaz +2
Vertical Federated Learning (VFL) has emerged as a critical paradigm for collaborative model training in privacy-sensitive domains such as finance and healthcare. However, most exi…
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
Imanol Miranda, Ander Salaberria, Eneko Agirre +1
Dual-encoder Vision-Language Models (VLMs) such as CLIP are often characterized as bag-of-words systems due to their poor performance on compositional benchmarks. We argue that thi…
Multimodal LLMs Do Not Compose Skills Optimally Across Modalities
Paula Ontalvilla, Aitor Ormazabal, Gorka Azkune
Skill composition is the ability to combine previously learned skills to solve new tasks. As neural networks acquire increasingly complex skills during their pretraining, it is not…
Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
Lukas Arana, Julen Etxaniz, Ander Salaberria +1
Current Multimodal Large Language Models exhibit very strong performance for several demanding tasks. While commercial MLLMs deliver acceptable performance in low-resource language…
Adding simple structure at inference improves Vision-Language Compositionality
Imanol Miranda, Ander Salaberria, Eneko Agirre +1
Dual encoder Vision-Language Models (VLM) such as CLIP are widely used for image-text retrieval tasks. However, those models struggle with compositionality, showing a bag-of-words-…