2 papers
cs.LG2025
Automatically Finding Rule-Based Neurons in OthelloGPT
Aditya Singh, Zihang Wen, Srujananjali Medicherla +2
OthelloGPT, a transformer trained to predict valid moves in Othello, provides an ideal testbed for interpretability research. The model is complex enough to exhibit rich computatio…
cs.CL2025
Concept Incongruence: An Exploration of Time and Death in Role Playing
Xiaoyan Bai, Ike Peng, Aditya Singh +1
Consider this prompt "Draw a unicorn with two horns". Should large language models (LLMs) recognize that a unicorn has only one horn by definition and ask users for clarifications,…