8 papers
Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models
Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
Do large language models contain domain-specific parametric shells: concentrated, causally necessary neuron populations whose removal selectively degrades a target domain while spa…
Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models
Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
We present a feedforward graph architecture in which heterogeneous frozen large language models serve as computational nodes, communicating through a shared continuous latent space…
Thinking in Different Spaces: Domain-Specific Latent Geometry Survives Cross-Architecture Translation
Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
We investigate whether independently trained language models converge to geometrically compatible latent representations, and whether this compatibility can be exploited to correct…
Say Anything but This: When Tokenizer Betrays Reasoning in LLMs
Navid Ayoobi, Marcus I Armstrong, Arjun Mukherjee
Large language models (LLMs) reason over discrete token ID sequences, yet modern subword tokenizers routinely produce non-unique encodings: multiple token ID sequences can detokeni…
Exposing Pink Slime Journalism: Linguistic Signatures and Robust Detection Against LLM-Generated Threats
Sadat Shahriar, Navid Ayoobi, Arjun Mukherjee +2
The local news landscape, a vital source of reliable information for 28 million Americans, faces a growing threat from Pink Slime Journalism, a low-quality, auto-generated articles…
The Erosion of LLM Signatures: Can We Still Distinguish Human and LLM-Generated Scientific Ideas After Iterative Paraphrasing?
Sadat Shahriar, Navid Ayoobi, Arjun Mukherjee
With the increasing reliance on LLMs as research agents, distinguishing between LLM and human-generated ideas has become crucial for understanding the cognitive nuances of LLMs' re…