6 papers
DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?
Young-Suk Lee, Ramon Fernandez Astudillo, Radu Florian
Deep research agents increasingly interleave web browsing with multi-step computation, yet existing benchmarks evaluate these capabilities in isolation, creating a blind spot in as…
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat +7
We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multiling…
Optimal Policy Minimum Bayesian Risk
Ramón Fernandez Astudillo, Md Arafat Sultan, Aashka Trivedi +4
Inference scaling helps LLMs solve complex reasoning problems through extended runtime computation. On top of long chain-of-thought (long-CoT) models, purely inference-time techniq…
Granite Embedding Models
Parul Awasthy, Aashka Trivedi, Yulong Li +19
We introduce the Granite Embedding models, a family of encoder-based embedding models designed for retrieval tasks, spanning dense-retrieval and sparse retrieval architectures, wit…
From Multiple-Choice to Extractive QA: A Case Study for English and Arabic
Teresa Lynn, Malik H. Altakrori, Samar Mohamed Magdy +11
The rapid evolution of Natural Language Processing (NLP) has favoured major languages such as English, leaving a significant gap for many others due to limited resources. This is e…
CLAPNQ: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems
Sara Rosenthal, Avirup Sil, Radu Florian +1
Retrieval Augmented Generation (RAG) has become a popular application for large language models. It is preferable that successful RAG systems provide accurate answers that are supp…