papers

Publications (12)

cs.CL2026

Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing

Raghavv Goel, Mukul Gagrani, Mingu Lee +1

Large Language Models (LLMs) possess latent multi-token prediction (MTP) abilities despite being trained only for next-token generation. We introduce ESP (Embedding-Space Probing),…

cs.AI2026

KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments

Junyoung Park, Dalton Jones, Matthew J Morse +3

We demonstrate that geometrically distinctive keys during LLM inference tend to have high attention scores. Based on the phenomenon we propose KeyDiff, a training-free KV cache evi…

cs.LG2024

Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement

Wonseok Jeon, Mukul Gagrani, Raghavv Goel +3

Speculative decoding is an inference-acceleration method for large language models (LLMs) where a small language model generates a draft-token sequence which is further verified by…

cs.LG2026

Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs

Sudhanshu Agrawal, Risheek Garrepalli, Raghavv Goel +3

Diffusion LLMs (dLLMs) have recently emerged as a powerful alternative to autoregressive LLMs (AR-LLMs) with the potential to operate at significantly higher token-generation rates…

eess.SY2022

Composite Adaptive Control for Time-varying Systems with Dual Adaptation

Raghavv Goel, Sayan Basu Roy

This paper proposes a composite adaptive control architecture using dual adaptation scheme for dynamical systems comprising time-varying uncertain parameters. While majority of the…

eess.IV2024

Motion Informed Needle Segmentation in Ultrasound Images

Raghavv Goel, Cecilia Morales, Manpreet Singh +3

Segmenting a moving needle in ultrasound images is challenging due to the presence of artifacts, noise, and needle occlusion. This task becomes even more demanding in scenarios whe…