9 citations · 16 across the 8 of their papers we have counts for
Showing 2025Show all
2 papers · 1 filter
cs.CL2025★ 9 cited
Reasoning Models Don't Always Say What They Think
Yanda Chen, Joe Benton, Ansh Radhakrishnan +12
Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effecti…
cs.CL2025
Counterfactual Simulatability of LLM Explanations for Generation Tasks
Marvin Limpijankit, Yanda Chen, Melanie Subbiah +2
LLMs can be unpredictable, as even slight alterations to the prompt can cause the output to change in unexpected ways. Thus, the ability of models to accurately explain their behav…