collaborators

10 papers

cs.LG2026

Understanding helpfulness and harmless tension in reward models

Eshaan Tanwar, Pepa Atanasova

Reward models are a key component of reinforcement learning from human feedback (RLHF), aligning language models toward both helpful and harmless behaviour. However, the internal m…

cs.CL2026

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

Jingyi Sun, Qianli Wang, Pepa Atanasova +2

Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated with metrics under two disjoint p…

cs.CL2026

Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models

Jingyi Sun, Pepa Atanasova, Sagnik Ray Choudhury +2

Context utilisation, the ability of Language Models (LMs) to incorporate relevant information from the provided context when generating responses, remains largely opaque to users,…

cs.CL2026

Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement

Sekh Mainul Islam, Pepa Atanasova, Isabelle Augenstein

Natural Language Explanations (NLEs) describe how Large Language Models (LLMs) make decisions by drawing on external Context Knowledge (CK) and Parametric Knowledge (PK). Understan…

cs.CL2026

Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations

Qianli Wang, Nils Feldhus, Pepa Atanasova +4

Quantization is widely used to accelerate inference and streamline the deployment of large language models (LLMs), yet its effects on self-explanations (SEs) remain unexplored. SEs…

cs.CL2025

Self-Critique and Refinement for Faithful Natural Language Explanations

Yingming Wang, Pepa Atanasova

With the rapid development of Large Language Models (LLMs), Natural Language Explanations (NLEs) have become increasingly important for understanding model predictions. However, th…