4 papers · 1 filter
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
Hakaze Cho, Haolin Yang, Yanshu Li +2
Existing works are dedicated to untangling atomized numerical components (features) from the hidden states of Large Language Models (LLMs). However, they typically rely on autoenco…
Mechanism of Task-oriented Information Removal in In-context Learning
Hakaze Cho, Haolin Yang, Gouki Minegishi +1
In-context Learning (ICL) is an emerging few-shot learning paradigm based on modern Language Models (LMs), yet its inner mechanism remains unclear. In this paper, we investigate th…
Pareto Optimal Algorithmic Recourse in Multi-cost Function
Wen-Ling Chen, Hong-Chang Huang, Kai-Hung Lin +2
In decision-making systems, algorithmic recourse aims to identify minimal-cost actions to alter an individual features, thereby obtaining a desired outcome. This empowers individua…
PXGen: A Post-hoc Explainable Method for Generative Models
Yen-Lung Huang, Ming-Hsi Weng, Hao-Tsung Yang
With the rapid growth of generative AI in numerous applications, explainable AI (XAI) plays a crucial role in ensuring the responsible development and deployment of generative AI t…