4 papers
Beyond Confidence: The Rhythms of Reasoning in Generative Models
Deyuan Liu, Zecheng Wang, Zhanyue Qin +3
Large Language Models (LLMs) exhibit impressive capabilities yet suffer from sensitivity to slight input context variations, hampering reliability. Conventional metrics like accura…
LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models
Zhanyue Qin, Yue Ding, Deyuan Liu +7
Nowadays, Large Language Models (LLMs) have attracted widespread attention due to their powerful performance. However, due to the unavoidable exposure to socially biased data durin…
Mitigating Gender Bias in Code Large Language Models via Model Editing
Zhanyue Qin, Haochuan Wang, Zecheng Wang +6
In recent years, with the maturation of large language model (LLM) technology and the emergence of high-quality programming code datasets, researchers have become increasingly conf…
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
Haochuan Wang, Xiachong Feng, Lei Li +4
The rapid advancement of large language models has accelerated their application in reasoning, with strategic reasoning drawing increasing attention. To evaluate the strategic reas…