1 paper
Mengyu Xu, Qiaoxin Yang, Qianqian Wang +3
Existing safety evaluations for large language models overlook whether responses preserve comparable medical information across different user phrasings of the same question. To ad…