Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes
A. Bochkov
Trainable input embedding tables are a standard component of modern language models. We ask whether they are actually necessary at the input interface. For a vocabulary of size …
cs.CL2025
Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations
A. Bochkov
Understanding the locus of semantic representation in large language models (LLMs) is crucial for interpretability and architectural innovation. The dominant paradigm posits that t…