2 citations · 3 across the 11 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
Guangzhi Xiong, Zhenghao He, Bohan Liu +2
Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs) by grounding outputs in retrieved evidence, but faithfulness failures, where generation…
cs.CL2021
Perturbing Inputs for Fragile Interpretations in Deep Natural Language Processing
Sanchit Sinha, Hanjie Chen, Arshdeep Sekhon +2
Interpretability methods like Integrated Gradient and LIME are popular choices for explaining natural language model predictions with relative word importance scores. These interpr…