activity
20232026
most citedAutomatic Textual Normalization for Hate Speech Detection

3 citations · 3 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

SAGE: Sustainable Agent-Guided Expert-tuning for Culturally Attuned Translation in Low-Resource Southeast Asia

Zhixiang Lu, Chong Zhang, Yulong Li +5

The vision of an inclusive World Wide Web is impeded by a severe linguistic divide, particularly for communities in low-resource regions of Southeast Asia. While large language mod…

cs.CL2026

DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs

Anh Thi-Hoang Nguyen, Khanh Quoc Tran, Tin Van Huynh +3

The reliability of large language models (LLMs) in production environments remains significantly constrained by their propensity to generate hallucinations -- fluent, plausible-sou…

cs.CL2025

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models

Nguyen Tien Dong, Minh-Anh Nguyen, Thanh Dat Hoang +6

The rapid advancement of large language models (LLMs) has enabled new possibilities for applying artificial intelligence within the legal domain. Nonetheless, the complexity, hiera…

cs.CL2025

ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization

Anh Thi-Hoang Nguyen, Dung Ha Nguyen, Kiet Van Nguyen

ViSoLex is an open-source system designed to address the unique challenges of lexical normalization for Vietnamese social media text. The platform provides two core services: Non-S…

cs.CL2024

A Weakly Supervised Data Labeling Framework for Machine Lexical Normalization in Vietnamese Social Media

Dung Ha Nguyen, Anh Thi Hoang Nguyen, Kiet Van Nguyen

This study introduces an innovative automatic labeling framework to address the challenges of lexical normalization in social media texts for low-resource languages like Vietnamese…

cs.CL2023★ 3 cited

Automatic Textual Normalization for Hate Speech Detection

Anh Thi-Hoang Nguyen, Dung Ha Nguyen, Nguyet Thi Nguyen +2

Social media data is a valuable resource for research, yet it contains a wide range of non-standard words (NSW). These irregularities hinder the effective operation of NLP tools. C…