◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Marco Cognetta

2 papers hereh-index 327 citations8 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author1

Across the 2 of 2 papers where every author was matched, so the position is known.

fields
  • cs.CL2
same name
  • Marco Cognetta — 1 paper, h 2

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.CLShow all

3 papers · 1 filter

cs.CL2025

Decoding-Free Sampling Strategies for LLM Marginalization

David Pohl, Marco Cognetta, Junyoung Lee +1

Modern language models operate on subword-tokenized text in order to make a trade-off between model size, inference speed, and vocabulary coverage. A side effect of this is that, d…

cs.CL2024

Tokenization as Finite-State Transduction

Marco Cognetta, Naoaki Okazaki

Tokenization is the first step in modern neural language model pipelines where an input text is converted to a sequence of subword tokens. We introduce from first principles a fini…

cs.CL2024

Distributional Properties of Subword Regularization

Marco Cognetta, Vilém Zouhar, Naoaki Okazaki

Subword regularization, used widely in NLP, improves model performance by reducing the dependency on exact tokenizations, augmenting the training corpus, and exposing the model to…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.