Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Pitfalls in Evaluating Interpretability Agents
Tal Haklay, Nikhil Prakash, Sana Pandey +5
Automated interpretability systems aim to reduce the need for human labor and scale analysis to increasingly large models and diverse tasks. Recent efforts toward this goal leverag…
cs.AI2023
Modeling Boundedly Rational Agents with Latent Inference Budgets
Athul Paul Jacob, Abhishek Gupta, Jacob Andreas
We study the problem of modeling a population of agents pursuing unknown goals subject to unknown computational constraints. In standard models of bounded rationality, sub-optimal…