4 papers · 1 filter
Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models
Junhao Liu, Haonan Yu, Zhenyu Yan +1
Post-hoc explanations provide transparency and are essential for guiding model optimization, such as prompt engineering and data sanitation. However, applying model-agnostic techni…
Beyond Attribution: Unified Concept-Level Explanations
Junhao Liu, Haonan Yu, Xin Zhang
There is an increasing need to integrate model-agnostic explanation techniques with concept-based approaches, as the former can explain models across different architectures while…
MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation
Haonan Yu, Junhao Liu, Xin Zhang
Anchors is a popular local model-agnostic explanation technique whose applicability is limited by its computational inefficiency. To address this limitation, we propose a memorizat…
ReX: A Framework for Incorporating Temporal Information in Model-Agnostic Local Explanation Techniques
Junhao Liu, Xin Zhang
Existing local model-agnostic explanation techniques are ineffective for machine learning models that consider inputs of variable lengths, as they do not consider temporal informat…