Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences
Ziyan Chen, Zhongzhu Zhou, Peilin Liu +1
Empirical scaling laws for autoregressive language models relate prediction loss to model size, data size, and optimization compute, but their theoretical origin is still poorly un…
cs.LG2026
Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization
Peilin Liu, Ding-Xuan Zhou
Transformer-based large models have demonstrated remarkable generalization abilities across different tasks by leveraging a context-aware attention module for in-context learning.…