Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
Hikaru Tsujimura, Arush Tagade
Large Language Models (LLMs) often display overconfidence, presenting information with unwarranted certainty in high-stakes contexts. We investigate the internal basis of this beha…
cs.LG2025
Benchmarking the Discovery Engine
Jack Foxabbott, Arush Tagade, Andrew Cusick +6
The Discovery Engine is a general purpose automated system for scientific discovery, which combines machine learning with state-of-the-art ML interpretability to enable rapid and r…
cs.LG2024
The SaTML '24 CNN Interpretability Competition: New Innovations for Concept-Level Interpretability
Stephen Casper, Jieun Yun, Joonhyuk Baek +13
Interpretability techniques are valuable for helping humans understand and oversee AI systems. The SaTML 2024 CNN Interpretability Competition solicited novel methods for studying…