Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
FastE: Readout-Triggered Token Compression for LLM Embedding Inference
Jinsong Shu, Jinyong Wen, Baokun Wang +4
In this study, we identify depth-dependent prefix redundancy in final-readout LLM embedding models, notably across representative backbones including Qwen3-Embedding and Qwen3-VL-E…
cs.AI2026
Token Economics for LLM Agents: A Dual-View Study from Computing and Economics
Yuxi Chen, Junming Chen, Chenyu He +9
As LLM agents evolve, tokens have emerged as the core economic primitives of Agentic AI. However, their exponential consumption introduces severe computational, collaborative, and…