Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
ETS: Efficient Tree Search for Inference-Time Scaling
Coleman Hooper, Sehoon Kim, Suhong Moon +7
Test-time compute scaling has emerged as a new axis along which to improve model accuracy, where additional computation is used at inference time to allow the model to think longer…
cs.LG2025
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
Rishabh Tiwari, Haocheng Xi, Aditya Tomar +7
Large Language Models (LLMs) are increasingly being deployed on edge devices for long-context settings, creating a growing need for fast and efficient long-context inference. In th…
cs.LG2024
Efficient and Scalable Estimation of Tool Representations in Vector Space
Suhong Moon, Siddharth Jha, Lutfi Eren Erdogan +4
Recent advancements in function calling and tool use have significantly enhanced the capabilities of large language models (LLMs) by enabling them to interact with external informa…