activity
20222025
most citedCausal Proxy Models for Concept-Based Model Explanations

6 citations · 12 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2025

Composing Policy Gradients and Prompt Optimization for Language Model Programs

Noah Ziems, Dilara Soylu, Lakshya A Agrawal +10

Group Relative Policy Optimization (GRPO) has proven to be an effective tool for post-training language models (LMs). However, AI systems are increasingly expressed as modular prog…

cs.CL2025

HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks

Jiuding Sun, Jing Huang, Sidharth Baskaran +4

Mechanistic interpretability has made great strides in identifying neural network features (e.g., directions in hidden activation space) that mediate concepts(e.g., the birth year…

cs.CL20245 cited

In-Context Learning for Extreme Multi-Label Classification

Karel D'Oosterlinck, Omar Khattab, François Remy +3

Multi-label classification problems with thousands of classes are hard to solve with in-context learning alone, as language models (LMs) might lack prior knowledge about the precis…

cs.CL2023

Flexible Model Interpretability through Natural Language Model Editing

Karel D'Oosterlinck, Thomas Demeester, Chris Develder +1

Model interpretability and model editing are crucial goals in the age of large language models. Interestingly, there exists a link between these two goals: if a method is able to s…

cs.CL2023

CAW-coref: Conjunction-Aware Word-level Coreference Resolution

Karel D'Oosterlinck, Semere Kiros Bitew, Brandon Papineau +3

State-of-the-art coreference resolutions systems depend on multiple LLM calls per document and are thus prohibitively expensive for many use cases (e.g., information extraction wit…

cs.CL2023

Rigorously Assessing Natural Language Explanations of Neurons

Jing Huang, Atticus Geiger, Karel D'Oosterlinck +2

Natural language is an appealing medium for explaining how large language models process and store information, but evaluating the faithfulness of such explanations is challenging.…