2 papers
cs.LG2026
Flash-GMM: A Memory-Efficient Kernel for Scalable Soft Clustering
Gal Bloch, Ariel Gera, Matan Orbach +2
We present \textbf{Flash-GMM}, a fused Triton kernel for efficient computation of Gaussian Mixture Models (GMMs) over large-scale data in a single GPU pass. By eliminating the need…
cs.CL2025
An Analysis of Hyper-Parameter Optimization Methods for Retrieval Augmented Generation
Matan Orbach, Ohad Eytan, Benjamin Sznajder +12
Optimizing Retrieval-Augmented Generation (RAG) configurations for specific tasks is a complex and resource-intensive challenge. Motivated by this challenge, frameworks for RAG hyp…