Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Concept-Level Explainability for Auditing & Steering LLM Responses
Kenza Amara, Rita Sevastjanova, Mennatallah El-Assady
As large language models (LLMs) become widely deployed, concerns about their safety and alignment grow. An approach to steer LLM behavior, such as mitigating biases or defending ag…
cs.CL2024
SyntaxShap: Syntax-aware Explainability Method for Text Generation
Kenza Amara, Rita Sevastjanova, Mennatallah El-Assady
To harness the power of large language models in safety-critical domains, we need to ensure the explainability of their predictions. However, despite the significant attention to m…
cs.CL2024
Challenges and Opportunities in Text Generation Explainability
Kenza Amara, Rita Sevastjanova, Mennatallah El-Assady
The necessity for interpretability in natural language processing (NLP) has risen alongside the growing prominence of large language models. Among the myriad tasks within NLP, text…