Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
LaCache: Robust Semantic Caching for LLM Serving
Jiacheng Liang, Yuhui Wang, Tanqiu Jiang +1
Semantic caching, which reuses responses to semantically similar requests via their embeddings, has seen growing adoption in LLM serving, offering faster responses and reduced cost…
cs.AI2026
Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language Models
Ting Wang, Yuanjie Shi, Yan Yan +1
Large language models (LLMs) increasingly perform multi-step reasoning, where intermediate claims form implicit directed acyclic graphs whose node correctness is structurally condi…
cs.AI2026
Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
Yuhui Wang, Changjiang Li, Guangke Chen +2
Large reasoning models (LRMs) exhibit unprecedented capabilities in solving complex problems through Chain-of-Thought (CoT) reasoning. However, recent studies reveal that their fin…