5 papers
Momentum for Reasoning: Dense Intrinsic Signals in Policy Optimization
Hao Chen, Zhanming Shen, Liyao Li +8
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for eliciting long-chain reasoning in large language models. However, existing methods base…
FLaG: Fine-Grained Latent Grouping for Hallucination Detection
Wentao Ye, Liyao Li, Zhiqing Xiao +6
Hallucinations in large language models (LLMs) arise from heterogeneous failure mechanisms, making reliable detection difficult for any single global uncertainty score. In this wor…
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
Hao Chen, Qi Zhang, Liyao Li +7
Adapting Large Language Models (LLMs) to specialized domains typically incurs high data and computational overhead. While prior efficiency efforts have largely treated data selecti…
Table as a Modality for Large Language Models
Liyao Li, Chao Ye, Wentao Ye +9
To migrate the remarkable successes of Large Language Models (LLMs), the community has made numerous efforts to generalize them to the table reasoning tasks for the widely deployed…
TableGPT-R1: Advancing Tabular Reasoning Through Reinforcement Learning
Saisai Yang, Qingyi Huang, Jing Yuan +13
Tabular data serves as the backbone of modern data analysis and scientific research. While Large Language Models (LLMs) fine-tuned via Supervised Fine-Tuning (SFT) have significant…