◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Dan John Velasco

9 papers hereh-index 6212 citations14 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author1
  • first author3
  • middle author3
  • last author1

Across the 8 of 9 papers where every author was matched, so the position is known.

fields
  • cs.CL7
  • cs.CV2

identity via Semantic Scholar / OpenAlex

activity
20202025
most citedSEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages

2 citations · 2 across the 6 of their papers we have counts for

collaborators
Showing 2025 · cs.CLShow all

3 papers · 2 filters

cs.CL2025

Rethinking the Role of Text Complexity in Language Model Pretraining

Dan John Velasco, Matthew Theodore Roque

Improving pretraining data quality and size is known to boost downstream performance, but the role of text complexity--how hard a text is to read--remains less explored. We reduce…

cs.CL2025

Beyond Repetition: Text Simplification and Curriculum Learning for Data-Constrained Pretraining

Matthew Theodore Roque, Dan John Velasco

Most studies on language model pretraining focus on large datasets, leaving open questions about optimization in data-constrained settings. In such settings, the effects of trainin…

cs.CL2025

Scaling, Simplification, and Adaptation: Lessons from Pretraining on Machine-Translated Text

Dan John Velasco, Matthew Theodore Roque

Most languages lack sufficient data for large-scale monolingual pretraining, creating a "data wall." Multilingual pretraining helps but is limited by language imbalance and the "cu…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.