Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Luxical: High-Speed Lexical-Dense Text Embeddings
DatologyAI, :, Luke Merrick +31
Frontier language model quality increasingly hinges on our ability to organize web-scale text corpora for training. Today's dominant tools trade off speed and flexibility: lexical…
cs.CL2024
Brevity is the soul of wit: Pruning long files for code generation
Aaditya K. Singh, Yu Yang, Kushal Tirumala +2
Data curation is commonly considered a "secret-sauce" for LLM training, with higher quality data usually leading to better LLM performance. Given the scale of internet-scraped corp…