information retrieval 1multi-agent collaboration 1open-domain search 1state management 1tool integration 1
From the 1 of 5 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Flatter Tokens are More Valuable for Speculative Draft Model Training
Jiaming Fan, Daming Cao, Xiangzhong Luo +3
Speculative Decoding (SD) is a key technique for accelerating Large Language Model (LLM) inference, but it typically requires training a draft model on a large dataset. We approach…
cs.CL2025
Fast Large Language Model Collaborative Decoding via Speculation
Jiale Fu, Yuchu Jiang, Junkai Chen +3
Large Language Model (LLM) collaborative decoding techniques improve output quality by combining the outputs of multiple models at each generation step, but they incur high computa…