2 papers
cs.CL2026
Retrieval Heads are Dynamic
Yuping Lin, Zitao Li, Yue Xing +6
Recent studies have identified "retrieval heads" in Large Language Models (LLMs) responsible for extracting information from input contexts. However, prior works largely rely on st…
cs.LG2024
Why Fine-grained Labels in Pretraining Benefit Generalization?
Guan Zhe Hong, Yin Cui, Ariel Fuxman +2
Recent studies show that pretraining a deep neural network with fine-grained labeled data, followed by fine-tuning on coarse-labeled data for downstream tasks, often yields better…