4 papers
P2S: Probabilistic Process Supervision for General-Domain Reasoning Question Answering
Wenlin Zhong, Chengyuan Liu, Yiquan Wu +5
While reinforcement learning with verifiable rewards (RLVR) has advanced LLM reasoning in structured domains like mathematics and programming, its application to general-domain rea…
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
Weikang Yuan, Kaisong Song, Zhuoren Jiang +6
Legal consultation is essential for safeguarding individual rights and ensuring access to justice, yet remains costly and inaccessible to many individuals due to the shortage of pr…
AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios
Yuting Huang, Meitong Guo, Yiquan Wu +6
Recent advances in LegalAI have primarily focused on individual case judgment analysis, often overlooking the critical appellate process within the judicial system. Appeals serve a…
Towards Stepwise Domain Knowledge-Driven Reasoning Optimization and Reflection Improvement
Chengyuan Liu, Shihang Wang, Lizhi Qing +7
Recently, stepwise supervision on Chain of Thoughts (CoTs) presents an enhancement on the logical reasoning tasks such as coding and math, with the help of Monte Carlo Tree Search…