Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
Omar Mahmoud, Ali Khalil, Buddhika Laknath Semage +2
Hallucination in large language models (LLMs) has been widely studied in recent years, with progress in both detection and mitigation aimed at improving truthfulness. Yet, a critic…
cs.CL2025
Improving Multilingual Language Models by Aligning Representations through Steering
Omar Mahmoud, Buddhika Laknath Semage, Thommen George Karimpanal +1
This paper investigates how Large Language Models (LLMs) represent non-English tokens -- a question that remains underexplored despite recent progress. We propose a lightweight int…
cs.CL2024
Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
Aly M. Kassem, Omar Mahmoud, Niloofar Mireshghallah +5
In this paper, we introduce a black-box prompt optimization method that uses an attacker LLM agent to uncover higher levels of memorization in a victim agent, compared to what is r…