4 citations · 5 across the 4 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
Yuxin Xiao, Chaoqun Wan, Yonggang Zhang +5
As the development and application of Large Language Models (LLMs) continue to advance rapidly, enhancing their trustworthiness and aligning them with human preferences has become…
cs.CL2024★ 1 cited
Interpreting and Improving Large Language Models in Arithmetic Calculation
Wei Zhang, Chaoqun Wan, Yonggang Zhang +4
Large language models (LLMs) have demonstrated remarkable potential across numerous applications and have shown an emergent ability to tackle complex reasoning tasks, such as mathe…
cs.CL2024
From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning
Wei Chen, Zhen Huang, Liang Xie +9
Large Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue. When challenged by users, LLMs tend t…