2 papers
cs.LG2026
Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them
Kevin Zhou, Lisa Alazraki, Kris Cao +1
Pre-training data mixtures are commonly tuned by running small-scale experiments and extrapolating to the target training budget. When high-quality data is scarce and must be repea…
cs.CL2025
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
Matthieu Meeus, Igor Shilov, Shubham Jain +3
Whether LLMs memorize their training data and what this means, from measuring privacy leakage to detecting copyright violations, has become a rapidly growing area of research. In t…