Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Beware of the Batch Size: Hyperparameter Bias in Evaluating LoRA
Sangyoon Lee, Jaeho Lee
Low-rank adaptation (LoRA) is a standard approach for fine-tuning large language models, yet its many variants report conflicting empirical gains, often on the same benchmarks. We…
cs.LG2025
eMamba: Efficient Acceleration Framework for Mamba Models in Edge Computing
Jiyong Kim, Jaeho Lee, Jiahao Lin +4
State Space Model (SSM)-based machine learning architectures have recently gained significant attention for processing sequential data. Mamba, a recent sequence-to-sequence SSM, of…