3 papers
cs.CL2026
Language models struggle with compartmentalization
Thomas Vincent Howe, David Wingate
In the training data used by large language models (LLMs), the same latent concept is often presented in multiple distinct ways: the same facts appear in English and Swahili; many…
cs.LG2024
Features that Make a Difference: Leveraging Gradients for Improved Dictionary Learning
Jeffrey Olmo, Jared Wilson, Max Forsey +3
Sparse Autoencoders (SAEs) are a promising approach for extracting neural network representations by learning a sparse and overcomplete decomposition of the network's internal acti…
cs.HC2023
AI Chat Assistants can Improve Conversations about Divisive Topics
Lisa P. Argyle, Ethan Busby, Joshua Gubler +4
A rapidly increasing amount of human conversation occurs online. But divisiveness and conflict can fester in text-based interactions on social media platforms, in messaging apps, a…