◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Chenglei Si

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL3
same name
  • Chenglei Si — 4 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedBetween words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

106 citations · 106 across the 3 of their papers we have counts for

collaborators

3 papers

cs.CL2023

Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations

Chenglei Si, Dan Friedman, Nitish Joshi +3

In-context learning (ICL) is an important paradigm for adapting large language models (LLMs) to new tasks, but the generalization behavior of ICL remains poorly understood. We inve…

cs.CL2023

READIN: A Chinese Multi-Task Benchmark with Realistic and Diverse Input Noises

Chenglei Si, Zhengyan Zhang, Yingfa Chen +3

For many real-world applications, the user-generated inputs usually contain various noises due to speech recognition errors caused by linguistic variations1 or typographical errors…

cs.CL2021★ 106 cited

Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky +8

What are the units of text that we want to model? From bytes to multi-word expressions, text can be analyzed and generated at many granularities. Until recently, most natural langu…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.