Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
Matthew Finlayson, Ilia Kulikov, Daniel M. Bikel +3
Large language models (LLMs) often struggle with knowledge intensive NLP tasks, such as answering "Who won the latest World Cup?" because the knowledge they learn during training m…
cs.CL2024
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts
Ganesh Jawahar, Haichuan Yang, Yunyang Xiong +10
Weight-sharing supernets are crucial for performance estimation in cutting-edge neural architecture search (NAS) frameworks. Despite their ability to generate diverse subnetworks w…