activity
20242026
collaborators

6 papers

cs.CL2026

TAKE: Trajectory-Aware Knowledge Estimation for Text Dataset Distillation

Tri-Nhan Vo, Dang Nguyen, Sunil Gupta

Large-scale text corpora have become a quiet bottleneck in modern NLP, not just in storage, but in the accumulated cost of training, fine-tuning, and continual learning. We propose…

cs.CR2026

Taint-Based Code Slicing for LLMs-based Malicious NPM Package Detection

Dang-Khoa Nguyen, Gia-Thang Ho, Quang-Minh Pham +5

Software supply chain attacks on the npm ecosystem have grown increasingly sophisticated, exploiting obfuscation and complex logic to evade detection. Large Language Models (LLMs)…

cs.CV2026

Improving Diversity in Black-box Few-shot Knowledge Distillation

Tri-Nhan Vo, Dang Nguyen, Kien Do +1

Knowledge distillation (KD) is a well-known technique to effectively compress a large network (teacher) to a smaller network (student) with little sacrifice in performance. However…

cs.LG2026

High-dimensional Level Set Estimation with Trust Regions and Double Acquisition Functions

Giang Ngo, Dat Phan Trong, Dang Nguyen +1

Level set estimation (LSE) classifies whether an unknown function's value exceeds a specified threshold for given inputs, a fundamental problem in many real-world applications. In…

cs.LG2025

Causal-Aware Generative Adversarial Networks with Reinforcement Learning

Tu Anh Hoang Nguyen, Dang Nguyen, Tri-Nhan Vo +2

The utility of tabular data for tasks ranging from model training to large-scale data analysis is often constrained by privacy concerns or regulatory hurdles. While existing data g…

cs.CV2024

Few-shot Algorithm Assurance

Dang Nguyen, Sunil Gupta

In image classification tasks, deep learning models are vulnerable to image distortion. For successful deployment, it is important to identify distortion levels under which the mod…