activity
20232026
most citedRethinking Interpretability in the Era of Large Language Models

43 citations · 60 across the 18 of their papers we have counts for

collaborators
Showing 2024Show all

9 papers · 1 filter

cs.CL2024

Interpretable Next-token Prediction via the Generalized Induction Head

Eunji Kim, Sriya Mantena, Weiwei Yang +3

While large transformer models excel in predictive performance, their lack of interpretability restricts their usefulness in high-stakes domains. To remedy this, we propose the Gen…

cs.CL2024

Vector-ICL: In-context Learning with Continuous Vector Representations

Yufan Zhuang, Chandan Singh, Liyuan Liu +2

Large language models (LLMs) have shown remarkable in-context learning (ICL) capabilities on textual data. We explore whether these capabilities can be extended to continuous vecto…

cs.CL2024

Generative causal testing to bridge data-driven models and scientific theories in language neuroscience

Richard Antonello, Chandan Singh, Shailee Jain +5

Representations from large language models are highly effective at predicting BOLD fMRI responses to language stimuli. However, these representations are largely opaque: it is uncl…

cs.CL2024

Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering

Qingru Zhang, Xiaodong Yu, Chandan Singh +6

Large language models (LLMs) have demonstrated remarkable performance across various real-world tasks. However, they often struggle to fully comprehend and effectively utilize thei…

cs.CL2024

Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts

Zeliang Zhang, Xiaodong Liu, Hao Cheng +2

By increasing model parameters but activating them sparsely when performing a task, the use of Mixture-of-Experts (MoE) architecture significantly improves the performance of Large…

cs.CL2024★ 2 cited

Crafting Interpretable Embeddings by Asking LLMs Questions

Vinamra Benara, Chandan Singh, John X. Morris +4

Large language models (LLMs) have rapidly improved text embeddings for a growing array of natural-language processing tasks. However, their opaqueness and proliferation into scient…