activity
20202025
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

Rethinking the Role of Text Complexity in Language Model Pretraining

Dan John Velasco, Matthew Theodore Roque

Improving pretraining data quality and size is known to boost downstream performance, but the role of text complexity--how hard a text is to read--remains less explored. We reduce…

cs.CL2025

Beyond Repetition: Text Simplification and Curriculum Learning for Data-Constrained Pretraining

Matthew Theodore Roque, Dan John Velasco

Most studies on language model pretraining focus on large datasets, leaving open questions about optimization in data-constrained settings. In such settings, the effects of trainin…

cs.CL2025

Scaling, Simplification, and Adaptation: Lessons from Pretraining on Machine-Translated Text

Dan John Velasco, Matthew Theodore Roque

Most languages lack sufficient data for large-scale monolingual pretraining, creating a "data wall." Multilingual pretraining helps but is limited by language imbalance and the "cu…

cs.CL2020

Pagsusuri ng RNN-based Transfer Learning Technique sa Low-Resource Language

Dan John Velasco

Low-resource languages such as Filipino suffer from data scarcity which makes it challenging to develop NLP applications for Filipino language. The use of Transfer Learning (TL) te…

cs.CL2020

Exploiting News Article Structure for Automatic Corpus Generation of Entailment Datasets

Jan Christian Blaise Cruz, Jose Kristian Resabal, James Lin +2

Transformers represent the state-of-the-art in Natural Language Processing (NLP) in recent years, proving effective even in tasks done in low-resource languages. While pretrained t…