6 papers
Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models
Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
Do large language models contain domain-specific parametric shells: concentrated, causally necessary neuron populations whose removal selectively degrades a target domain while spa…
When Flat Minima Fail: Characterizing INT4 Quantization Collapse After FP32 Convergence
Marcus Armstrong
Post-training quantization (PTQ) assumes that a well-converged model is a quantization-ready model. We show this assumption fails in a structured, measurable, and previously unchar…
Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models
Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
We present a feedforward graph architecture in which heterogeneous frozen large language models serve as computational nodes, communicating through a shared continuous latent space…
Investigating the Fundamental Limit: A Feasibility Study of Hybrid-Neural Archival
Marcus Armstrong, ZiWei Qiu, Huy Q. Vo +1
Large Language Models (LLMs) possess a theoretical capability to model information density far beyond the limits of classical statistical methods (e.g., Lempel-Ziv). However, utili…
Thinking in Different Spaces: Domain-Specific Latent Geometry Survives Cross-Architecture Translation
Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
We investigate whether independently trained language models converge to geometrically compatible latent representations, and whether this compatibility can be exploited to correct…
Say Anything but This: When Tokenizer Betrays Reasoning in LLMs
Navid Ayoobi, Marcus I Armstrong, Arjun Mukherjee
Large language models (LLMs) reason over discrete token ID sequences, yet modern subword tokenizers routinely produce non-unique encodings: multiple token ID sequences can detokeni…