6 papers
Transparent Trade-offs between Properties of Explanations
Hiwot Belay Tadesse, Alihan Hüyük, Yaniv Yacoby +2
When explaining black-box machine learning models, it's often important for explanations to have certain desirable properties. Most existing methods `encourage' desirable propertie…
Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models
Salma Abdel Magid, Weiwei Pan, Simon Warchol +4
Text-to-image (T2I) models are increasingly used in impactful real-life applications. As such, there is a growing need to audit these models to ensure that they generate desirable,…
Inverse Reinforcement Learning with Multiple Planning Horizons
Jiayu Yao, Weiwei Pan, Finale Doshi-Velez +1
In this work, we study an inverse reinforcement learning (IRL) problem where the experts are planning under a shared reward function but with different, unknown planning horizons.…
A Sim2Real Approach for Identifying Task-Relevant Properties in Interpretable Machine Learning
Eura Nofshin, Esther Brown, Brian Lim +2
Explanations of an AI's function can assist human decision-makers, but the most useful explanation depends on the decision's context, referred to as the downstream task. User studi…
What Makes a Good Explanation?: A Harmonized View of Properties of Explanations
Zixi Chen, Varshini Subhash, Marton Havasi +2
Interpretability provides a means for humans to verify aspects of machine learning (ML) models and empower human+ML teaming in situations where the task cannot be fully automated.…
Towards Model-Agnostic Posterior Approximation for Fast and Accurate Variational Autoencoders
Yaniv Yacoby, Weiwei Pan, Finale Doshi-Velez
Inference for Variational Autoencoders (VAEs) consists of learning two models: (1) a generative model, which transforms a simple distribution over a latent space into the distribut…