9 citations · 12 across the 3 of their papers we have counts for
3 papers
cs.CL2025★ 9 cited
Reasoning Models Don't Always Say What They Think
Yanda Chen, Joe Benton, Ansh Radhakrishnan +12
Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effecti…
cs.LG2024★ 2 cited
Sabotage Evaluations for Frontier Models
Joe Benton, Misha Wagner, Eric Christiansen +13
Sufficiently capable models could subvert human oversight and decision-making in important contexts. For example, in the context of AI development, models could covertly sabotage e…
cs.LG2023★ 1 cited
Measuring Feature Sparsity in Language Models
Mingyang Deng, Lucas Tao, Joe Benton
Recent works have proposed that activations in language models can be modelled as sparse linear combinations of vectors corresponding to features of input text. Under this assumpti…