3 papers
cs.CL2026
BCMT: Blockwise Causal Memory Transformer
Rachid Arezki
Transformer architectures rely on dense self-attention to model long-range dependencies, but this mechanism exhibits quadratic complexity with respect to sequence length. We introd…
cs.LG2026
Soft Contamination Means Benchmarks Test Shallow Generalization
Ari Spiesberger, Juan J. Vazquez, Nicky Pochinkov +4
If LLM training data is polluted with benchmark test data, then benchmark performance gives biased estimates of out-of-distribution (OOD) generalization. Typical decontamination fi…
cs.CL2025
AI-AI Bias: large language models favor communications generated by large language models
Walter Laurito, Benjamin Davis, Peli Grietzer +3
Are large language models (LLMs) biased in favor of communications produced by LLMs, leading to possible antihuman discrimination? Using a classical experimental design inspired by…