8 papers
GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno +1
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabi…
Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding
Karim Radouane, Jose G Moreno, Lynda Tamine
Understanding concepts is fundamental to generalization. Despite their impressive performance on a wide range of tasks, Large Language Models (LLMs) still struggle with genuine con…
Improving Ad-hoc Search Effectiveness for Conversational Information Retrieval via Model Merging
Ahmed Rayane Kebir, Jose G. Moreno, Lynda Tamine
Conversational information retrieval is challenging since it requires the consideration of the conversation history which potentially gives rise to topic shifts and coreference res…
ReToP: Learning to Rewrite Electronic Health Records for Clinical Prediction
Jesus Lovon-Melgarejo, Jose G. Moreno, Christine Damase-Michel +1
Electronic Health Records (EHRs) provide crucial information for clinical decision-making. However, their high-dimensionality, heterogeneity, and sparsity make clinical prediction…
Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens
Lucas Albarede, Jose Moreno, Lynda Tamine +1
Despite their impressive performances, Large Language Models (LLMs) remain prone to hallucination, which critically undermines their trustworthiness. While most of the previous wor…
Revisiting the MIMIC-IV Benchmark: Experiments Using Language Models for Electronic Health Records
Jesus Lovon, Thouria Ben-Haddi, Jules Di Scala +2
The lack of standardized evaluation benchmarks in the medical domain for text inputs can be a barrier to widely adopting and leveraging the potential of natural language models for…