1 paper · 1 filter
Ali Mahdavi, Azaseh Zamanifar, Amirfarhad Farhadi +1
Long-prompt inference remains expensive because prefill attention scales quadratically with sequence length. We propose Spectral-LSH, a training-free prompt compression method that…