4 papers
P2S: Probabilistic Process Supervision for General-Domain Reasoning Question Answering
Wenlin Zhong, Chengyuan Liu, Yiquan Wu +5
While reinforcement learning with verifiable rewards (RLVR) has advanced LLM reasoning in structured domains like mathematics and programming, its application to general-domain rea…
AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios
Yuting Huang, Meitong Guo, Yiquan Wu +6
Recent advances in LegalAI have primarily focused on individual case judgment analysis, often overlooking the critical appellate process within the judicial system. Appeals serve a…
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
Weikang Yuan, Kaisong Song, Zhuoren Jiang +6
Legal consultation is essential for safeguarding individual rights and ensuring access to justice, yet remains costly and inaccessible to many individuals due to the shortage of pr…
Towards Stepwise Domain Knowledge-Driven Reasoning Optimization and Reflection Improvement
Chengyuan Liu, Shihang Wang, Lizhi Qing +7
Recently, stepwise supervision on Chain of Thoughts (CoTs) presents an enhancement on the logical reasoning tasks such as coding and math, with the help of Monte Carlo Tree Search…