5 papers
MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning
Haotian Wang, Lian Yan, Xingzhi Yao +4
In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated using a tolerance-based reward.…
ProSpec RL: Plan Ahead, then Execute
Liangliang Liu, Yi Guan, BoRan Wang +5
Imagining potential outcomes of actions before execution helps agents make more informed decisions, a prospective thinking ability fundamental to human cognition. However, mainstre…
KCS: Diversify Multi-hop Question Generation with Knowledge Composition Sampling
Yangfan Wang, Jie Liu, Chen Tang +2
Multi-hop question answering faces substantial challenges due to data sparsity, which increases the likelihood of language models learning spurious patterns. To address this issue,…
KA2L: A Knowledge-Aware Active Learning Framework for LLMs
Haoxuan Yin, Bojian Liu, Chen Tang +3
Fine-tuning large language models (LLMs) with high-quality knowledge has been shown to enhance their performance effectively. However, there is a paucity of research on the depth o…
AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models
Lian Yan, Haotian Wang, Chen Tang +5
In the agricultural domain, the deployment of large language models (LLMs) is hindered by the lack of training data and evaluation benchmarks. To mitigate this issue, we propose Ag…