4 papers
SHARD: Safe and Helpful Alignment via Self-Reframing Distillation
Viswonathan Manoranjan, Amogh Gupta, Anvesh Rao Vijjini +2
Large language models often struggle with sensitive prompts. They may refuse outright, provide generic safety boilerplate, or fail to address the user's legitimate informational ne…
AI-Mediated Explainable Regulation for Justice
Thomas Hofweber, Andreas Sudmann, Evangelos Pournaras
Present practice of deciding on regulation faces numerous problems that make adopted regulations static, unexplained, unduly influenced by powerful interest groups, and stained wit…
Are language models rational? The case of coherence norms and belief revision
Thomas Hofweber, Peter Hase, Elias Stengel-Eskin +1
Do norms of rationality apply to machine learning models, in particular language models? In this paper we investigate this question by focusing on a special subset of rational norm…
The Black Tuesday Attack: how to crash the stock market with adversarial examples to financial forecasting models
Thomas Hofweber, Jefrey Bergl, Ian Reyes +1
We investigate and defend the possibility of causing a stock market crash via small manipulations of individual stock values that together realize an adversarial example to financi…