6 papers
LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks
Yifan Chen, Haitao Li, Yiran Hu +6
As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential. These tasks…
Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation
Yongjie Wang, Xinyue Zhang, Kunhong Yao +4
Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference. Such age…
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
Yige Xu, Yongjie Wang, Zizhuo Wu +3
Reasoning in vision-language models (VLMs) has recently attracted significant attention due to its broad applicability across diverse downstream tasks. However, it remains unclear…
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
Weikang Yuan, Kaisong Song, Zhuoren Jiang +6
Legal consultation is essential for safeguarding individual rights and ensuring access to justice, yet remains costly and inaccessible to many individuals due to the shortage of pr…
Towards Stepwise Domain Knowledge-Driven Reasoning Optimization and Reflection Improvement
Chengyuan Liu, Shihang Wang, Lizhi Qing +7
Recently, stepwise supervision on Chain of Thoughts (CoTs) presents an enhancement on the logical reasoning tasks such as coding and math, with the help of Monte Carlo Tree Search…
Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator
Chengyuan Liu, Shihang Wang, Lizhi Qing +4
Domain Large Language Models (LLMs) are developed for domain-specific tasks based on general LLMs. But it still requires professional knowledge to facilitate the expertise for some…