activity
20172025
most citedHSD Shared Task in VLSP Campaign 2019:Hate Speech Detection for Social Good

20 citations · 51 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2025

MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources

Huu Nguyen, Victor May, Harsh Raj +14

We present MixtureVitae, an open-access pretraining corpus built to minimize legal risk while providing strong downstream performance. MixtureVitae follows a permissive-first, risk…

cs.CL2025

The Efficiency of Pre-training with Objective Masking in Pseudo Labeling for Semi-Supervised Text Classification

Arezoo Hatefi, Xuan-Son Vu, Monowar Bhuyan +1

We extend and study a semi-supervised model for text classification proposed earlier by Hatefi et al. for classification tasks in which document classes are described by a small nu…

cs.CL2024

Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code

Taishi Nakamura, Mayank Mishra, Simone Tedeschi +42

Pretrained language models are an integral part of AI applications, but their high computational cost for training limits accessibility. Initiatives such as Bloom and StarCoder aim…

cs.CL2023

Grandma Karl is 27 years old -- research agenda for pseudonymization of research data

Elena Volodina, Simon Dobnik, Therese Lindström Tiedemann +1

Accessibility of research data is critical for advances in many research fields, but textual data often cannot be shared due to the personal and sensitive information which it cont…

cs.CL202020 cited

HSD Shared Task in VLSP Campaign 2019:Hate Speech Detection for Social Good

Xuan-Son Vu, Thanh Vu, Mai-Vu Tran +2

The paper describes the organisation of the "HateSpeech Detection" (HSD) task at the VLSP workshop 2019 on detecting the fine-grained presence of hate speech in Vietnamese textual…

cs.CL20194 cited

Generic Multilayer Network Data Analysis with the Fusion of Content and Structure

Xuan-Son Vu, Abhishek Santra, Sharma Chakravarthy +1

Multi-feature data analysis (e.g., on Facebook, LinkedIn) is challenging especially if one wants to do it efficiently and retain the flexibility by choosing features of interest fo…