1 citations · 1 across the 1 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
The Power of Architecture: Deep Dive into Transformer Architectures for Long-Term Time Series Forecasting
Lefei Shen, Mouxiang Chen, Han Fu +5
Transformer-based models have recently become dominant in Long-term Time Series Forecasting (LTSF), yet the variations in their architecture, such as encoder-only, encoder-decoder,…
cs.LG2025
Parallel Scaling Law for Language Models
Mouxiang Chen, Binyuan Hui, Zeyu Cui +5
It is commonly believed that scaling language models should commit a significant space or time cost, by increasing the parameters (parameter scaling) or output tokens (inference-ti…