Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Full Glyph Images Beat Token Embeddings: A Controlled Study for Transformers
Shuyang Xiang, Hao Guan
Modern language models generally represent text as sequences of discrete token embeddings, an assumption deeply rooted in current practice but rarely questioned. We challenge this…
cs.CV2026
Hot-Start Chinese Language Modeling:Visual Glyphs Accelerate Sample-Efficient Learning
Shuyang Xiang, Hao Guan
In this work, we study whether rendering Chinese characters as visual glyph images, rather than discrete token IDs as mainstream LLMs do, providing an inductive bias for character-…