3 papers
cs.LG2026
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders
Shunchang Liu, Xin Chen, Belen Martin Urcelay +1
Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit preference instability, producing contr…
cs.CL2025
The Biased Oracle: Assessing LLMs' Understandability and Empathy in Medical Diagnoses
Jianzhou Yao, Shunchang Liu, Guillaume Drui +3
Large language models (LLMs) show promise for supporting clinicians in diagnostic communication by generating explanations and guidance for patients. Yet their ability to produce o…
cs.CV2025
CopyJudge: Automated Copyright Infringement Identification and Mitigation in Text-to-Image Diffusion Models
Shunchang Liu, Zhuan Shi, Lingjuan Lyu +2
Assessing whether AI-generated images are substantially similar to source works is a crucial step in resolving copyright disputes. In this paper, we propose CopyJudge, a novel auto…