10 papers
Understanding helpfulness and harmless tension in reward models
Eshaan Tanwar, Pepa Atanasova
Reward models are a key component of reinforcement learning from human feedback (RLHF), aligning language models toward both helpful and harmless behaviour. However, the internal m…
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
Jingyi Sun, Qianli Wang, Pepa Atanasova +2
Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated with metrics under two disjoint p…
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
Jingyi Sun, Pepa Atanasova, Sagnik Ray Choudhury +2
Context utilisation, the ability of Language Models (LMs) to incorporate relevant information from the provided context when generating responses, remains largely opaque to users,…
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
Sekh Mainul Islam, Pepa Atanasova, Isabelle Augenstein
Natural Language Explanations (NLEs) describe how Large Language Models (LLMs) make decisions by drawing on external Context Knowledge (CK) and Parametric Knowledge (PK). Understan…
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
Qianli Wang, Nils Feldhus, Pepa Atanasova +4
Quantization is widely used to accelerate inference and streamline the deployment of large language models (LLMs), yet its effects on self-explanations (SEs) remain unexplored. SEs…
Self-Critique and Refinement for Faithful Natural Language Explanations
Yingming Wang, Pepa Atanasova
With the rapid development of Large Language Models (LLMs), Natural Language Explanations (NLEs) have become increasingly important for understanding model predictions. However, th…