2 papers
cs.CR2026
When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs
Jiaming Cheng, Subhransu Das, Rajiv Ramnath
Multi-agent LLM systems relay key-value caches instead of text and credit their gains to exchanged "latent thoughts". That credit is a claim about which example's cache is relayed,…
cs.AI2026
Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
Subhransu Das, Jiaming Cheng, Arnav Kumar +6
Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not deploy well. W…