2 papers
cs.CL2025
Length-MAX Tokenizer for Language Models
Dong Dong, Weijie Su
We introduce a new tokenizer for language models that minimizes the average tokens per character, thereby reducing the number of tokens needed to represent text during training and…
cs.HC2024
Human-LLM Collaborative Construction of a Cantonese Emotion Lexicon
Yusong Zhang, Dong Dong, Chi-tim Hung +3
Large Language Models (LLMs) have demonstrated remarkable capabilities in language understanding and generation. Advanced utilization of the knowledge embedded in LLMs for automate…