1 citations · 2 across the 6 of their papers we have counts for
3 papers · 1 filter
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
Swapnaja Achintalwar, Ioana Baldini, Djallel Bouneffouf +16
The alignment of large language models is usually done by model providers to add or control behaviors that are common or universally understood across use cases and contexts. In co…
Unveiling Safety Vulnerabilities of Large Language Models
George Kour, Marcel Zalmanovici, Naama Zwerdling +5
As large language models become more prevalent, their possible harmful or inappropriate responses are a cause for concern. This paper introduces a unique dataset containing adversa…
Predicting Question-Answering Performance of Large Language Models through Semantic Consistency
Ella Rabinovich, Samuel Ackerman, Orna Raz +2
Semantic consistency of a language model is broadly defined as the model's ability to produce semantically-equivalent outputs, given semantically-equivalent inputs. We address the…