activity
20142024
most citedProceedings of the 2016 ICML Workshop on Human Interpretability in Machine Learning (WHI 2016)

24 citations · 49 across the 13 of their papers we have counts for

collaborators

15 papers

cs.CR20241 cited

Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI

Ambrish Rawat, Stefan Schoepf, Giulio Zizzo +10

As generative AI, particularly large language models (LLMs), become increasingly integrated into production applications, new attack surfaces and vulnerabilities emerge and put a f…

cs.CL2024

Value Alignment from Unstructured Text

Inkit Padhi, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri +3

Aligning large language models (LLMs) to value systems has emerged as a significant area of research within the fields of AI and NLP. Currently, this alignment process relies on th…

cs.CY2024

When Trust is Zero Sum: Automation Threat to Epistemic Agency

Emmie Malone, Saleh Afroogh, Jason DCruz +1

AI researchers and ethicists have long worried about the threat that automation poses to human dignity, autonomy, and to the sense of personal value that is tied to work. Typically…

cs.AI20241 cited

Contextual Moral Value Alignment Through Context-Based Aggregation

Pierre Dognin, Jesus Rios, Ronny Luss +7

Developing value-aligned AI agents is a complex undertaking and an ongoing challenge in the field of AI. Specifically within the domain of Large Language Models (LLMs), the capabil…

cs.LG2024

A resource-constrained stochastic scheduling algorithm for homeless street outreach and gleaning edible food

Conor M. Artman, Aditya Mate, Ezinne Nwankwo +10

We developed a common algorithmic solution addressing the problem of resource-constrained outreach encountered by social change organizations with different missions and operations…

cs.CL2024

Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations

Swapnaja Achintalwar, Ioana Baldini, Djallel Bouneffouf +16

The alignment of large language models is usually done by model providers to add or control behaviors that are common or universally understood across use cases and contexts. In co…