Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
Sungkyun Kim, Jaemin Kim, Dogyung Yoon +3
LLMs have low GPU efficiency and high latency due to autoregressive decoding. Speculative decoding (SD) mitigates this using a small draft model to speculatively generate multiple…
cs.CL2024
Benchmarks Underestimate the Readiness of Multi-lingual Dialogue Agents
Andrew H. Lee, Sina J. Semnani, Galo Castillo-López +16
Creating multilingual task-oriented dialogue (TOD) agents is challenging due to the high cost of training data acquisition. Following the research trend of improving training data…