3 papers
cs.AI2026
SAE-StatSteer: Statistical Consensus Feature Selection for Optimization-Free Activation Steering of Large Language Models
Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy +3
Activation steering adds a residual-stream direction at inference time, providing lightweight behavioral control without fine-tuning. Sparse autoencoders (SAEs) can make such inter…
cs.HC2026
Can Generative AI help people navigate Radical Moral Disagreements? The CONSIDER prototype
William Hohnen-Ford, Sarah Chen, Kathryn B. Francis +3
Radical Moral Disagreements (RMDs) are highly polarising topics that are increasingly censored in everyday life, with growing evidence suggesting that this polarisation carries mea…
cs.CL2025
Rethinking Word Similarity: Semantic Similarity through Classification Confusion
Kaitlyn Zhou, Haishan Gao, Sarah Chen +3
Word similarity has many applications to social science and cultural analytics tasks like measuring meaning change over time and making sense of contested terms. Yet traditional si…