2 papers
cs.AI2025
Uncovering the Computational Ingredients of Human-Like Representations in LLMs
Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee +1
The ability to translate diverse patterns of inputs into structured patterns of behavior has been thought to rest on both humans' and machines' ability to learn robust representati…
cs.AI2025
Evaluating Steering Techniques using Human Similarity Judgments
Zach Studdiford, Timothy T. Rogers, Siddharth Suresh +1
Current evaluations of Large Language Model (LLM) steering techniques focus on task-specific performance, overlooking how well steered representations align with human cognition. U…