activity
20242026
collaborators

6 papers

cs.CL2026

LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks

Yifan Chen, Haitao Li, Yiran Hu +6

As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential. These tasks…

cs.CR2026

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation

Yongjie Wang, Xinyue Zhang, Kunhong Yao +4

Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference. Such age…

cs.CV2026

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap

Yige Xu, Yongjie Wang, Zizhuo Wu +3

Reasoning in vision-language models (VLMs) has recently attracted significant attention due to its broad applicability across diverse downstream tasks. However, it remains unclear…

cs.CL2025

LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation

Weikang Yuan, Kaisong Song, Zhuoren Jiang +6

Legal consultation is essential for safeguarding individual rights and ensuring access to justice, yet remains costly and inaccessible to many individuals due to the shortage of pr…

cs.AI2025

Towards Stepwise Domain Knowledge-Driven Reasoning Optimization and Reflection Improvement

Chengyuan Liu, Shihang Wang, Lizhi Qing +7

Recently, stepwise supervision on Chain of Thoughts (CoTs) presents an enhancement on the logical reasoning tasks such as coding and math, with the help of Monte Carlo Tree Search…

cs.CL2024

Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator

Chengyuan Liu, Shihang Wang, Lizhi Qing +4

Domain Large Language Models (LLMs) are developed for domain-specific tasks based on general LLMs. But it still requires professional knowledge to facilitate the expertise for some…