2 papers
cs.AI2025
Learnable Game-theoretic Policy Optimization for Data-centric Self-explanation Rationalization
Yunxiao Zhao, Zhiqiang Wang, Xingtong Yu +3
Rationalization, a data-centric framework, aims to build self-explanatory models to explain the prediction outcome by generating a subset of human-intelligible pieces of the input…
cs.CL2025
Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
Yunxiao Zhao, Hao Xu, Zhiqiang Wang +3
Pre-trained Language Models (PLMs) are trained on large amounts of unlabeled data, yet they exhibit remarkable reasoning skills. However, the trustworthiness challenges posed by th…