Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
A Systematic Study of Training-Free Methods for Trustworthy Large Language Models
Wai Man Si, Mingjie Li, Michael Backes +1
As Large Language Models (LLMs) receive increasing attention and are being deployed across various domains, their potential risks, including generating harmful or biased content, p…
cs.CL2026
Triviality Corrected Endogenous Reward
Xinda Wang, Zhengxu Hou, Yangshijie Zhang +6
Reinforcement learning for open-ended text generation is constrained by the lack of verifiable rewards, necessitating reliance on judge models that require either annotated data or…
cs.CL2026
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
Mingjie Li, Wai Man Si, Michael Backes +2
Despite the impressive performance of general-purpose large language models (LLMs), they often require fine-tuning or post-training to excel at specific tasks. For instance, large…