3 papers
cs.CV2026
Full Glyph Images Beat Token Embeddings: A Controlled Study for Transformers
Shuyang Xiang, Hao Guan
Modern language models generally represent text as sequences of discrete token embeddings, an assumption deeply rooted in current practice but rarely questioned. We challenge this…
cs.CV2026
Hot-Start Chinese Language Modeling:Visual Glyphs Accelerate Sample-Efficient Learning
Shuyang Xiang, Hao Guan
In this work, we study whether rendering Chinese characters as visual glyph images, rather than discrete token IDs as mainstream LLMs do, providing an inductive bias for character-…
cs.LG2025
PAODING: A High-fidelity Data-free Pruning Toolkit for Debloating Pre-trained Neural Networks
Mark Huasong Meng, Hao Guan, Liuhuo Wan +3
We present PAODING, a toolkit to debloat pretrained neural network models through the lens of data-free pruning. To preserve the model fidelity, PAODING adopts an iterative process…