3 papers
cs.CR2026
ToxScreen: Detecting Whether an LLM Has Been Poisoned
Anthony Hughes, Nicole Xing, Collin Francel +2
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behav…
cs.LG2026
SMIXAE: Towards Unsupervised Manifold Discovery in Language Models
Collin Francel
Sparse autoencoders (SAEs) have been used widely to decompose and interpret neural network activations, especially those of transformer language models. One key issue with SAEs is…
cs.CL2025
Language Models for Adult Service Website Text Analysis
Nickolas Freeman, Thanh Nguyen, Gregory Bott +2
Sex trafficking refers to the use of force, fraud, or coercion to compel an individual to perform in commercial sex acts against their will. Adult service websites (ASWs) have and…