Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression
Noel Elias, Homa Esfahanizadeh, Kaan Kale +2
Large language models have drastically changed the prospects of AI by introducing technologies for more complex natural language processing. However, current methodologies to train…
cs.CL2024
OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset
Allen Roush, Yusuf Shabazz, Arvind Balaji +7
We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community. This dataset includes over 3.…
cs.CL2024
TexShape: Information Theoretic Sentence Embedding for Language Models
Kaan Kale, Homa Esfahanizadeh, Noel Elias +3
With the exponential growth in data volume and the emergence of data-intensive applications, particularly in the field of machine learning, concerns related to resource utilization…