2 papers
cs.CL2026
Compressing Sequences in the Latent Embedding Space: -Token Merging for Large Language Models
Zihao Xu, John Harvill, Ziwei Fan +3
Large Language Models (LLMs) incur significant computational and memory costs when processing long prompts, as full self-attention scales quadratically with input length. Token com…
cs.CL2025
Lossless Token Sequence Compression via Meta-Tokens
John Harvill, Ziwei Fan, Hao Wang +4
Existing work on prompt compression for Large Language Models (LLM) focuses on lossy methods that try to maximize the retention of semantic information that is relevant to downstre…