8 citations · 17 across the 8 of their papers we have counts for
4 papers · 1 filter
Pitfalls in Evaluating Interpretability Agents
Tal Haklay, Nikhil Prakash, Sana Pandey +5
Automated interpretability systems aim to reduce the need for human labor and scale analysis to increasingly large models and diverse tasks. Recent efforts toward this goal leverag…
Agents of Chaos
Natalie Shapira, Chris Wendler, Avery Yen +35
We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord acc…
Discovering Variable Binding Circuitry with Desiderata
Xander Davies, Max Nadeau, Nikhil Prakash +2
Recent work has shown that computation in language models may be human-understandable, with successful efforts to localize and intervene on both single-unit features and input-outp…
Conceptualization and Framework of Hybrid Intelligence Systems
Nikhil Prakash, Kory W. Mathewson
As artificial intelligence (AI) systems are getting ubiquitous within our society, issues related to its fairness, accountability, and transparency are increasing rapidly. As a res…