4 papers
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization
Jorge L. Ruiz Williams
KV-cache quantizers usually optimize storage-space reconstruction, even though attention reads keys through logits and values through attention-weighted readout. We argue that pers…
The Condensate Theorem: Transformers are O(n), Not
Jorge L. Ruiz Williams
We present the Condensate Theorem: attention sparsity is a learned topological property, not an architectural constraint. Through empirical analysis of trained language models, we…
Warp-Cortex: An Asynchronous, Memory-Efficient Architecture for Million-Agent Cognitive Scaling on Consumer Hardware
Jorge L. Ruiz Williams
Current multi-agent Large Language Model (LLM) frameworks suffer from linear memory scaling, rendering "System 2" parallel reasoning impractical on consumer hardware. We present Wa…
Fast Witness Persistence for MRI Volumes via Hybrid Landmarking
Jorge Leonardo Ruiz Williams
We introduce a scalable witness-based persistent homology pipeline for full-brain MRI volumes that couples density-aware landmark selection with a GPU-ready witness filtration. Can…