2 papers
cs.LG2026
Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context
Alagappan Valliappan
Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in M…
cs.SD2021
Memory-efficient Speech Recognition on Smart Devices
Ganesh Venkatesh, Alagappan Valliappan, Jay Mahadeokar +4
Recurrent transducer models have emerged as a promising solution for speech recognition on the current and next generation smart devices. The transducer models provide competitive…