3 citations · 5 across the 4 of their papers we have counts for
5 papers
Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation
Leen AlQadi, Ahmed Alzubaidi, Mohammed Alyafeai +6
We present QIMMA, a quality-assured Arabic LLM leaderboard that places systematic benchmark validation at its core. Rather than aggregating existing resources as-is, QIMMA applies…
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
Mouadh Yagoubi, Yasser Dahou, Billel Mokeddem +12
Existing benchmarks have proven effective for assessing the performance of fully trained large language models. However, we find striking differences in the early training stages o…
Falcon2-11B Technical Report
Quentin Malartic, Nilabhra Roy Chowdhury, Ruxandra Cojocaru +14
We introduce Falcon2-11B, a foundation model trained on over five trillion tokens, and its multimodal counterpart, Falcon2-11B-vlm, which is a vision-to-text model. We report our f…
Deep Retrieval-Based Dialogue Systems: A Short Review
Basma El Amel Boussaha, Nicolas Hernandez, Christine Jacquin +1
Building dialogue systems that naturally converse with humans is being an attractive and an active research domain. Multiple systems are being designed everyday and several dataset…
Thread Reconstruction in Conversational Data using Neural Coherence Models
Dat Tien Nguyen, Shafiq Joty, Basma El Amel Boussaha +1
Discussion forums are an important source of information. They are often used to answer specific questions a user might have and to discover more about a topic of interest. Discuss…