2 papers
cs.CR2026
Privacy-Aware Split Inference with Speculative Decoding for Large Language Models over Wide-Area Networks
Michael Cunningham
We present a practical system for privacy-aware large language model (LLM) inference that splits a transformer between a trusted local GPU and an untrusted cloud GPU, communicating…
cs.CL2025
Contextual Compression Encoding for Large Language Models: A Novel Framework for Multi-Layered Parameter Space Pruning
Barnaby Schmitt, Alistair Grosvenor, Matthias Cunningham +3
Context-aware compression techniques have gained increasing attention as model sizes continue to grow, introducing computational bottlenecks that hinder efficient deployment. A str…