6 papers
TAKE: Trajectory-Aware Knowledge Estimation for Text Dataset Distillation
Tri-Nhan Vo, Dang Nguyen, Sunil Gupta
Large-scale text corpora have become a quiet bottleneck in modern NLP, not just in storage, but in the accumulated cost of training, fine-tuning, and continual learning. We propose…
Taint-Based Code Slicing for LLMs-based Malicious NPM Package Detection
Dang-Khoa Nguyen, Gia-Thang Ho, Quang-Minh Pham +5
Software supply chain attacks on the npm ecosystem have grown increasingly sophisticated, exploiting obfuscation and complex logic to evade detection. Large Language Models (LLMs)…
Improving Diversity in Black-box Few-shot Knowledge Distillation
Tri-Nhan Vo, Dang Nguyen, Kien Do +1
Knowledge distillation (KD) is a well-known technique to effectively compress a large network (teacher) to a smaller network (student) with little sacrifice in performance. However…
High-dimensional Level Set Estimation with Trust Regions and Double Acquisition Functions
Giang Ngo, Dat Phan Trong, Dang Nguyen +1
Level set estimation (LSE) classifies whether an unknown function's value exceeds a specified threshold for given inputs, a fundamental problem in many real-world applications. In…
Causal-Aware Generative Adversarial Networks with Reinforcement Learning
Tu Anh Hoang Nguyen, Dang Nguyen, Tri-Nhan Vo +2
The utility of tabular data for tasks ranging from model training to large-scale data analysis is often constrained by privacy concerns or regulatory hurdles. While existing data g…
Few-shot Algorithm Assurance
Dang Nguyen, Sunil Gupta
In image classification tasks, deep learning models are vulnerable to image distortion. For successful deployment, it is important to identify distortion levels under which the mod…