4 papers
SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation
Yimeng Zhang, Yingying Zhuang, Ziyi Wang +12
Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However,…
SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents
Ziyi Wang, Yuxuan Lu, Yimeng Zhang +8
Language model agents are increasingly effective in solving realistic tasks through multi-turn tool use. However, training reliable tool-using agents remains challenging in practic…
Meta-RTL: Reinforcement-Based Meta-Transfer Learning for Low-Resource Commonsense Reasoning
Yu Fu, Jie He, Yifan Yang +2
Meta learning has been widely used to exploit rich-resource source tasks to improve the performance of low-resource target tasks. Unfortunately, most existing meta learning approac…
LFED: A Literary Fiction Evaluation Dataset for Large Language Models
Linhao Yu, Qun Liu, Deyi Xiong
The rapid evolution of large language models (LLMs) has ushered in the need for comprehensive assessments of their performance across various dimensions. In this paper, we propose…