Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
SCIT: Testing Causal Cache Carriers in Latent Chain-of-Thought Models
Yi Ding, Lijun Huang, Menglin Yang
Latent chain-of-thought models move intermediate reasoning from emitted text into continuous states, improving compactness but hiding the causal object. We introduce SCIT, the Suff…
cs.CL2026
FCPRAG: Fusion-Controller Parametric Retrieval-Augmented Generation for Stable Multi-Passage LoRA Injection
Jinchang Zhu, Jindong Li, Yi Ding +5
Parametric retrieval-augmented generation (PRAG) injects retrieved evidence into a large language model (LLM) through passage-specific LoRA adapters, reducing reliance on long in-c…
cs.CL2026
PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference
Niqi Lyu, Pengtao Shi, Wei Qiu +4
Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on diffic…