2 papers
cs.CL2026
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression
Noel Elias, Homa Esfahanizadeh, Kaan Kale +2
Large language models have drastically changed the prospects of AI by introducing technologies for more complex natural language processing. However, current methodologies to train…
cs.CL2024
TexShape: Information Theoretic Sentence Embedding for Language Models
Kaan Kale, Homa Esfahanizadeh, Noel Elias +3
With the exponential growth in data volume and the emergence of data-intensive applications, particularly in the field of machine learning, concerns related to resource utilization…