4.5k citations · 5.3k across the 47 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Accelerating LLM Inference via Vector Index Based Output Embeddings
Martin Loretz, Sepp Hochreiter
Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with large multilingual vocabularies. W…
cs.CL2026
Unlocking the Working Memory of Large Language Models for Latent Reasoning
Lukas Aichberger, Sepp Hochreiter
To improve the reasoning capabilities of large language models, test-time compute is typically scaled by generating intermediate tokens before the final answer. However, this coupl…