2 papers
cs.CL2026
Aligning What LLMs Do and Say: Towards Self-Consistent Explanations
Sahar Admoni, Ofra Amir, Assaf Hallak +1
Large language models (LLMs) seem to offer an easy path to interpretability: just ask them to explain their answers. Yet the features driving an answer often differ from those emph…
cs.LG2026
From Actions to Words: Towards Abstractive-Textual Policy Summarization in RL
Sahar Admoni, Assaf Hallak, Yftah Ziser +2
Explaining reinforcement learning agents is challenging because policies emerge from complex reward structures and neural representations that are difficult for humans to interpret…