2 papers
cs.CL2025
Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a Time
Huihan Li, You Chen, Siyuan Wang +4
Large Language Models (LLMs) perform well on reasoning benchmarks but often fail when inputs alter slightly, raising concerns about the extent to which their success relies on memo…
cs.CL2024
LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models
Haitao Li, You Chen, Qingyao Ai +3
Large language models (LLMs) have made significant progress in natural language processing tasks and demonstrate considerable potential in the legal domain. However, legal applicat…