Showing 2024Show all
2 papers · 1 filter
cs.LG2024
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
Martin Pawelczyk, Lillian Sun, Zhenting Qi +2
The rapid proliferation of generative AI, especially large language models, has led to their integration into a variety of applications. A key phenomenon known as weak-to-strong ge…
cs.AI2024
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
Tessa Han, Aounon Kumar, Chirag Agarwal +1
As large language models (LLMs) develop increasingly sophisticated capabilities and find applications in medical settings, it becomes important to assess their medical safety due t…