20 citations · 51 across the 9 of their papers we have counts for
10 papers · 1 filter
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
Huu Nguyen, Victor May, Harsh Raj +14
We present MixtureVitae, an open-access pretraining corpus built to minimize legal risk while providing strong downstream performance. MixtureVitae follows a permissive-first, risk…
The Efficiency of Pre-training with Objective Masking in Pseudo Labeling for Semi-Supervised Text Classification
Arezoo Hatefi, Xuan-Son Vu, Monowar Bhuyan +1
We extend and study a semi-supervised model for text classification proposed earlier by Hatefi et al. for classification tasks in which document classes are described by a small nu…
Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code
Taishi Nakamura, Mayank Mishra, Simone Tedeschi +42
Pretrained language models are an integral part of AI applications, but their high computational cost for training limits accessibility. Initiatives such as Bloom and StarCoder aim…
Grandma Karl is 27 years old -- research agenda for pseudonymization of research data
Elena Volodina, Simon Dobnik, Therese Lindström Tiedemann +1
Accessibility of research data is critical for advances in many research fields, but textual data often cannot be shared due to the personal and sensitive information which it cont…
HSD Shared Task in VLSP Campaign 2019:Hate Speech Detection for Social Good
Xuan-Son Vu, Thanh Vu, Mai-Vu Tran +2
The paper describes the organisation of the "HateSpeech Detection" (HSD) task at the VLSP workshop 2019 on detecting the fine-grained presence of hate speech in Vietnamese textual…
Generic Multilayer Network Data Analysis with the Fusion of Content and Structure
Xuan-Son Vu, Abhishek Santra, Sharma Chakravarthy +1
Multi-feature data analysis (e.g., on Facebook, LinkedIn) is challenging especially if one wants to do it efficiently and retain the flexibility by choosing features of interest fo…