8 papers
AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models
Mohammad Zbeeb, Hasan Abed Al Kader Hammoud, Sina Mukalled +5
We present AraLingBench: a fully human annotated benchmark for evaluating the Arabic linguistic competence of large language models (LLMs). The benchmark spans five core categories…
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
Hasan Abed Al Kader Hammoud, Mohammad Zbeeb, Bernard Ghanem
We present Hala, a family of Arabic-centric instruction and translation models built with our translate-and-tune pipeline. We first compress a strong AREN teacher…
Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
Mohammad Zbeeb, Hasan Abed Al Kader Hammoud, Bernard Ghanem
Large language models often require costly optimization, such as reinforcement learning, to master complex reasoning tasks. This work demonstrates that reasoning ability, once lear…
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
Hasan Abed Al Kader Hammoud, Kumail Alhamoud, Abed Hammoud +3
Recent work on enhancing the reasoning abilities of large language models (LLMs) has introduced explicit length control as a means of constraining computational cost while preservi…
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
Harethah Abu Shairah, Hasan Abed Al Kader Hammoud, Bernard Ghanem +1
Large language models (LLMs) are typically aligned to refuse harmful instructions through safety fine-tuning. A recent attack, termed abliteration, identifies and suppresses the si…
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
Hasan Abed Al Kader Hammoud, Hani Itani, Bernard Ghanem
Large Language Models (LLMs) leverage step-by-step reasoning to solve complex problems. Standard evaluation practice involves generating a complete reasoning trace and assessing th…