papers
Publications (2)
cs.IR2024
Mixture-of-PageRanks: Replacing Long-Context with Real-Time, Sparse GraphRAG
Nicholas Alonso, Beren Millidge
Recent advances have extended the context window of frontier LLMs dramatically, from a few thousand tokens up to millions, enabling entire books and codebases to fit into context.…
cs.CL2026
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
Tomas Figliolia, Nicholas Alonso, Rishi Iyer +2
Multi-headed Attention's (MHA) quadratic compute and linearly growing KV-cache make long-context transformers expensive to train and serve. Prior works such as Grouped Query Attent…