Publications (20)
Business (mis)Use Cases of Generative AI
Stephanie Houde, Vera Liao, Jacquelyn Martino +5
Generative AI is a class of machine learning technology that learns to generate new data from training data. While deep fakes and media-and art-related generative AI breakthroughs…
Iterative Design of Gestures During Elicitation: Understanding the Role of Increased Production
Andreea Danielescu, David Piorkowski
Previous gesture elicitation studies have found that user proposals are influenced by legacy bias which may inhibit users from proposing gestures that are most appropriate for an i…
Experiences with Improving the Transparency of AI Models and Services
Michael Hind, Stephanie Houde, Jacquelyn Martino +4
AI models and services are used in a growing number of highstakes areas, resulting in a need for increased transparency. Consistent with this, several proposals for higher quality…
Detecting Egregious Conversations between Customers and Virtual Agents
Tommy Sandbank, Michal Shmueli-Scheuer, Jonathan Herzig +3
Virtual agents are becoming a prominent channel of interaction in customer service. Not all customer interactions are smooth, however, and some can become almost comically bad. In…
Can LLM Code Explanations Adapt to Diverse Problem-Solvers' Needs?
Andrew Anderson, David Piorkowski, Justin Weisz +2
Large language model (LLM) code explanations can support people in solving code-related problems, yet prior work has shown that people have diverse problem-solving styles. If expla…
Quantitative AI Risk Assessments: Opportunities and Challenges
David Piorkowski, Michael Hind, John Richards
Although AI systems are increasingly being leveraged to provide value to organizations, individuals, and society, significant attendant risks have been identified and have manifest…
A Methodology for Creating AI FactSheets
John Richards, David Piorkowski, Michael Hind +2
As AI models and services are used in a growing number of highstakes areas, a consensus is forming around the need for a clearer record of how these models and services are develop…
Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations
Swapnaja Achintalwar, Adriana Alvarado Garcia, Ateret Anaby-Tavor +35
Large language models (LLMs) are susceptible to a variety of risks, from non-faithful output to biased and toxic generations. Due to several limiting factors surrounding LLMs (trai…
Towards evaluating and eliciting high-quality documentation for intelligent systems
David Piorkowski, Daniel González, John Richards +1
A vital component of trust and transparency in intelligent systems built on machine learning and artificial intelligence is the development of clear, understandable documentation.…
AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources
Frank Bagehorn, Kristina Brimijoin, Elizabeth M. Daly +17
The rapid evolution of generative AI has expanded the breadth of risks associated with AI systems. While various taxonomies and frameworks exist to classify these risks, the lack o…
FactSheets: Increasing Trust in AI Services through Supplier's Declarations of Conformity
Matthew Arnold, Rachel K. E. Bellamy, Michael Hind +10
Accuracy is an important concern for suppliers of artificial intelligence (AI) services, but considerations beyond accuracy, such as safety (which includes fairness and explainabil…
How AI Developers Overcome Communication Challenges in a Multidisciplinary Team: A Case Study
David Piorkowski, Soya Park, April Yi Wang +3
The development of AI applications is a multidisciplinary effort, involving multiple roles collaborating with the AI developers, an umbrella term we use to include data scientists…
Developing a Risk Identification Framework for Foundation Model Uses
David Piorkowski, Michael Hind, John Richards +1
As foundation models grow in both popularity and capability, researchers have uncovered a variety of ways that the models can pose a risk to the model's owner, user, or others. Des…
An LLM's Attempts to Adapt to Diverse Software Engineers' Problem-Solving Styles: More Inclusive & Equitable?
Andrew Anderson, David Piorkowski, Margaret Burnett +1
Software engineers use code-fluent large language models (LLMs) to help explain unfamiliar code, yet LLM explanations are not adapted to engineers' diverse problem-solving needs. W…
The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products
Hao-Ping Lee, Jessica He, David Piorkowski +3
Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. However, these same characteristics can create or exacerbate produ…
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
Anna Sokol, Elizabeth Daly, Michael Hind +4
Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation method…
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
Erik Miehling, Manish Nagireddy, Prasanna Sattigeri +3
Modern language models, while sophisticated, exhibit some inherent shortcomings, particularly in conversational settings. We claim that many of the observed shortcomings can be att…
Supporting Annotators with Affordances for Efficiently Labeling Conversational Data
Austin Z. Henley, David Piorkowski
Without well-labeled ground truth data, machine learning-based systems would not be as ubiquitous as they are today, but these systems rely on substantial amounts of correctly labe…
Evaluating a Methodology for Increasing AI Transparency: A Case Study
David Piorkowski, John Richards, Michael Hind
In reaction to growing concerns about the potential harms of artificial intelligence (AI), societies have begun to demand more transparency about how AI models and systems are crea…
Facilitating Knowledge Sharing from Domain Experts to Data Scientists for Building NLP Models
Soya Park, April Wang, Ban Kawas +3
Data scientists face a steep learning curve in understanding a new domain for which they want to build machine learning (ML) models. While input from domain experts could offer val…