2 papers
cs.LG2026
Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
Kehao Zhang, Shangtong Gui, Sheng Yang +2
Long-context LLMs and Retrieval-Augmented Generation (RAG) systems process information passively, deferring state tracking, contradiction resolution, and evidence aggregation to qu…
cs.LG2024
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
Zhuofan Wen, Shangtong Gui, Yang Feng
Inference acceleration of large language models (LLMs) has been put forward in many application scenarios and speculative decoding has shown its advantage in addressing inference a…