2 papers
cs.CL2024
Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-context Models
Xinyu Liu, Runsong Zhao, Pengcheng Huang +5
Numerous recent works target to extend effective context length for language models and various methods, tasks and benchmarks exist to measure model's effective memorization length…
cs.CL2024
NDP: Next Distribution Prediction as a More Broad Target
Junhao Ruan, Abudukeyumu Abudula, Xinyu Liu +7
Large language models (LLMs) trained on next-token prediction (NTP) paradigm have demonstrated powerful capabilities. However, the existing NTP paradigm contains several limitation…