5 papers
LLM Abstention Can Be a Prompt Artifact, in Addition to Genuine Uncertainty
Zipeng Ling, Shuliang Liu, Yuehao Tang +7
Large Language Models (LLMs) are increasingly trained to abstain from answering questions they are unsure about. However, this ability is often misused: in real-world applications,…
Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis
Zipeng Ling, Shuliang Liu, Shenghong Fu +4
LLM reasoning traces suffer from complex flaws -- *Step Internal Flaws* (logical errors, hallucinations, etc.) and *Step-wise Flaws* (overthinking, underthinking), which vary by sa…
AppellateGen: A Benchmark for Appellate Legal Judgment Generation
Hongkun Yang, Lionel Z. Wang, Wei Fan +10
Legal judgment generation is a critical task in legal intelligence. However, existing research in legal judgment generation has predominantly focused on first-instance trials, rely…
Quantifying LLM Biases Across Instruction Boundary in Mixed Question Forms
Zipeng Ling, Shuliang Liu, Yuehao Tang +8
Large Language Models (LLMs) annotated datasets are widely used nowadays, however, large-scale annotations often show biases in low-quality datasets. For example, Multiple-Choice Q…
Recurrent Knowledge Identification and Fusion for Language Model Continual Learning
Yujie Feng, Xujia Wang, Zexin Lu +7
Continual learning (CL) is crucial for deploying large language models (LLMs) in dynamic real-world environments without costly retraining. While recent model ensemble and model me…