Publications (5)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
Ge Yan, Tuomas Oikarinen, Tsui-Wei +1
Neuron identification is a popular tool in mechanistic interpretability, aiming to uncover the human-interpretable concepts represented by individual neurons in deep networks. Whil…
ReflCtrl: Controlling LLM Reflection via Representation Engineering
Ge Yan, Chung-En Sun, Tsui-Wei +1
Large language models (LLMs) with Chain-of-Thought (CoT) reasoning have achieved strong performance across diverse tasks, including mathematics, coding, and general reasoning. A di…
The Computational Complexity of Finding Arithmetic Expressions With and Without Parentheses
Jayson Lynch, Yan, Weng
We show NP-completeness for various problems about the existence of arithmetic expression trees. When given a set of operations, inputs, and a target value does there exist an expr…
Towards Verifying Robustness of Neural Networks Against Semantic Perturbations
Jeet Mohapatra, Tsui-Wei, Weng +3
Verifying robustness of neural networks given a specified threat model is a fundamental yet challenging task. While current verification methods mainly focus on the -norm t…
Hidden Cost of Randomized Smoothing
Jeet Mohapatra, Ching-Yun Ko, Tsui-Wei +4
The fragility of modern machine learning models has drawn a considerable amount of attention from both academia and the public. While immense interests were in either crafting adve…