385 citations · 766 across the 28 of their papers we have counts for
Showing 2023 · cs.AIShow all
2 papers · 2 filters
cs.AI2023★ 3 cited
Testing Language Model Agents Safely in the Wild
Silen Naihin, David Atkinson, Marc Green +5
A prerequisite for safe autonomy-in-the-wild is safe testing-in-the-wild. Yet real-world autonomous tests face several unique safety challenges, both due to the possibility of caus…
cs.AI2023
Discovering Variable Binding Circuitry with Desiderata
Xander Davies, Max Nadeau, Nikhil Prakash +2
Recent work has shown that computation in language models may be human-understandable, with successful efforts to localize and intervene on both single-unit features and input-outp…