4 papers
CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization
Deep Shah, Sanket Badhe, Nehal Kathrotia +1
Large Language Models utilizing reasoning techniques improve task performance but incur significant latency and token costs due to verbose generation. Existing automatic prompt opt…
Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications
Sanket Badhe, Deep Shah, Nehal Kathrotia
Large language models (LLMs) are trained on web-scale corpora that exhibit steep power-law distributions, in which the distribution of knowledge is highly long-tailed, with most ap…
Taxonomy of the Retrieval System Framework: Pitfalls and Paradigms
Deep Shah, Sanket Badhe, Nehal Kathrotia
Designing an embedding retrieval system requires navigating a complex design space of conflicting trade-offs between efficiency and effectiveness. This work structures these decisi…
Collective Communication for 100k+ GPUs
Min Si, Pavan Balaji, Yongzhou Chen +36
The increasing scale of large language models (LLMs) necessitates highly efficient collective communication frameworks, particularly as training workloads extend to hundreds of tho…