2 papers
cs.LG2025
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
Rishabh Tiwari, Haocheng Xi, Aditya Tomar +7
Large Language Models (LLMs) are increasingly being deployed on edge devices for long-context settings, creating a growing need for fast and efficient long-context inference. In th…
cs.LG2024
Efficient and Scalable Estimation of Tool Representations in Vector Space
Suhong Moon, Siddharth Jha, Lutfi Eren Erdogan +4
Recent advancements in function calling and tool use have significantly enhanced the capabilities of large language models (LLMs) by enabling them to interact with external informa…