6 papers
Rethinking Language Model Scaling under Transferable Hypersphere Optimization
Liliang Ren, Yang Liu, Yelong Shen +1
Scaling laws for large language models depend critically on the optimizer and parameterization. Existing hyperparameter transfer laws are mainly developed for first-order optimizer…
CoT2-Meta: Budgeted Metacognitive Control for Test-Time Reasoning
Siyuan Ma, Bo Gao, Zikai Xiao +6
Recent test-time reasoning methods improve performance by generating more candidate chains or searching over larger reasoning trees, but they typically lack explicit control over w…
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
Hossam Amer, Maryam Dialameh, Hossein Rajabzadeh +3
Scaling training compute, measured in FLOPs, has long been shown to improve the accuracy of large language models, yet training remains resource-intensive. Prior work shows that in…
ETT: Expanding the Long Context Understanding Capability of LLMs at Test-Time
Kiarash Zahirnia, Zahra Golpayegani, Walid Ahmed +1
Transformer-based Language Models' computation and memory overhead increase quadratically as a function of sequence length. The quadratic cost poses challenges when employing LLMs…
Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection
Mohammad Mahdi Moradi, Hossam Amer, Sudhir Mudur +3
Learning to adapt pretrained language models to unlabeled, out-of-distribution data is a critical challenge, as models often falter on structurally novel reasoning tasks even while…
Balancing Computation Load and Representation Expressivity in Parallel Hybrid Neural Networks
Mohammad Mahdi Moradi, Walid Ahmed, Shuangyue Wen +3
Attention and State-Space Models (SSMs) when combined in a hybrid network in sequence or in parallel provide complementary strengths. In a hybrid sequential pipeline they alternate…