20 citations · 20 across the 3 of their papers we have counts for
3 papers
cs.CL2025
Understanding Refusal in Language Models with Sparse Autoencoders
Wei Jie Yeo, Nirmalendu Prakash, Clement Neo +3
Refusal is a key safety behavior in aligned language models, yet the internal mechanisms driving refusals remain opaque. In this work, we conduct a mechanistic study of refusal in…
cs.CL2024
Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models
Wei Jie Yeo, Ranjan Satapathy, Erik Cambria
Large Language Models (LLMs) are capable of generating persuasive Natural Language Explanations (NLEs) to justify their answers. However, the faithfulness of these explanations sho…
cs.AI2023★ 20 cited
A Comprehensive Review on Financial Explainable AI
Wei Jie Yeo, Wihan van der Heever, Rui Mao +3
The success of artificial intelligence (AI), and deep learning models in particular, has led to their widespread adoption across various industries due to their ability to process…