73 citations · 107 across the 13 of their papers we have counts for
4 papers · 1 filter
Signature Activation: A Sparse Signal View for Holistic Saliency
Jose Roberto Tello Ayala, Akl C. Fahed, Weiwei Pan +4
The adoption of machine learning in healthcare calls for model transparency and explainability. In this work, we introduce Signature Activation, a saliency method that generates ho…
Why do universal adversarial attacks work on large language models?: Geometry might be the answer
Varshini Subhash, Anna Bialas, Weiwei Pan +1
Transformer based large language models with emergent capabilities are becoming increasingly ubiquitous in society. However, the task of understanding and interpreting their intern…
SAP-sLDA: An Interpretable Interface for Exploring Unstructured Text
Charumathi Badrinath, Weiwei Pan, Finale Doshi-Velez
A common way to explore text corpora is through low-dimensional projections of the documents, where one hopes that thematically similar documents will be clustered together in the…
The Unintended Consequences of Discount Regularization: Improving Regularization in Certainty Equivalence Reinforcement Learning
Sarah Rathnam, Sonali Parbhoo, Weiwei Pan +2
Discount regularization, using a shorter planning horizon when calculating the optimal policy, is a popular choice to restrict planning to a less complex set of policies when estim…