Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching
Bo Lv, Jingbo Sun
Optimizing the trade-off among predictive performance and computational cost is a central focus in the deployment of Large Language Models (LLMs). Current routing methods primarily…
cs.AI2026
OneLatent: Single-Token Compression for Visual Latent Reasoning
Bo Lv, Yasheng Sun, Junjie Wang +1
Chain-of-thought (CoT) prompting improves reasoning but often increases inference cost by one to two orders of magnitude. To address these challenges, we present \textbf{OneLatent}…