Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models
Xilin Gong, Shu Yang, Zehua Cao +2
Large Language Models (LLMs) have demonstrated strong capabilities for hidden representation interpretation through Patchscopes, a framework that uses LLMs themselves to generate h…
cs.CL2025
Investigating CoT Monitorability in Large Reasoning Models
Shu Yang, Junchao Wu, Xilin Gong +4
Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex tasks by engaging in extended reasoning before producing final answers. Beyond improving abilities…