From the 1 of 6 linked papers with an AI index.
6 papers
ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models
Ruxi Gu, Zhenliang Zhang, Wei Wang
The paper introduces ForgetBench, a benchmark for measuring how large language models retain or forget factual and relational knowledge when they are continuously edited over time.
SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States
Zhenliang Zhang, Wenqing Wang, Yong Hu +4
Long-Text Understanding (LTU) at million-token scale requires balancing reasoning fidelity with computational efficiency. Frontier long-context LLMs can process millions of token c…
SCOPE: Intrinsic Semantic Space Control for Mitigating Copyright Infringement in LLMs
Zhenliang Zhang, Xinyu Hu, Xiaojun Wan
Large language models sometimes inadvertently reproduce passages that are copyrighted, exposing downstream applications to legal risk. Most existing studies for inference-time defe…
JointCQ: Improving Factual Hallucination Detection with Joint Claim and Query Generation
Fan Xu, Huixuan Zhang, Zhenliang Zhang +2
Current large language models (LLMs) often suffer from hallucination issues, i,e, generating content that appears factual but is actually unreliable. A typical hallucination detect…
Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
Zhenliang Zhang, Junzhe Zhang, Xinyu Hu +2
Large language models (LLMs) have achieved remarkable success in various tasks, yet they remain vulnerable to faithfulness hallucinations, where the output does not align with the…
ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs
Zhenliang Zhang, Xinyu Hu, Huixuan Zhang +2
Large language models (LLMs) excel at various natural language processing tasks, but their tendency to generate hallucinations undermines their reliability. Existing hallucination…