Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Top-: Not All Logits Are You Need
Chenxia Tang, Jianchun Liu, Hongli Xu +1
Large language models (LLMs) typically employ greedy decoding or low-temperature sampling for reasoning tasks, reflecting a perceived trade-off between diversity and accuracy. We c…
cs.LG2024
Heterogeneous Learning Rate Scheduling for Neural Architecture Search on Long-Tailed Datasets
Chenxia Tang
In this paper, we attempt to address the challenge of applying Neural Architecture Search (NAS) algorithms, specifically the Differentiable Architecture Search (DARTS), to long-tai…