◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Omri Uzan

6 papers hereh-index 4149 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author4
  • middle author2

Across the 6 of 6 papers where every author was matched, so the position is known.

fields
  • cs.CL6

identity via Semantic Scholar / OpenAlex

activity
20242026
most citedEvaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge

6 citations · 8 across the 6 of their papers we have counts for

collaborators
Showing 2024Show all

3 papers · 1 filter

cs.CL2024★ 6 cited

Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge

Khuyagbaatar Batsuren, Ekaterina Vylomova, Verna Dankers +4

The popular subword tokenizers of current language models, such as Byte-Pair Encoding (BPE), are known not to respect morpheme boundaries, which affects the downstream performance…

cs.CL2024

Greed is All You Need: An Evaluation of Tokenizer Inference Methods

Omri Uzan, Craig W. Schmidt, Chris Tanner +1

While subword tokenizers such as BPE and WordPiece are typically used to build vocabularies for NLP models, the method of decoding text into a sequence of tokens from these vocabul…

cs.CL2024★ 2 cited

Tokenization Is More Than Compression

Craig W. Schmidt, Varshini Reddy, Haoran Zhang +4

Tokenization is a foundational step in natural language processing (NLP) tasks, bridging raw text and language models. Existing tokenization approaches like Byte-Pair Encoding (BPE…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.