Publications (31)
Multilevel Regression and Poststratification Interface: An Application to Track Community-level COVID-19 Viral Transmission
Yajuan Si, Toan Tran, Jonah Gabry +2
We present a novel Bayesian workflow for multilevel regression and poststratification (MRP), introducing extensions to time-varying data and granular geography and publicly availab…
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
Viet Nguyen, Anh Nguyen, Trung Dao +4
The escalating demand for real-time image synthesis has driven significant advancements in one-step diffusion models, which inherently offer expedited generation speeds compared to…
Differentially Private Synthetic Data via APIs 4: Tabular Data
Toan Tran, Arturs Backurs, Zinan Lin +3
This paper investigates the problem of generating synthetic tabular data with differential privacy (DP) guarantees, enabling data sharing in sensitive domains. Despite extensive st…
Learning Compositional Sparse Gaussian Processes with a Shrinkage Prior
Anh Tong, Toan Tran, Hung Bui +1
Choosing a proper set of kernel functions is an important problem in learning Gaussian Process (GP) models since each kernel structure has different model complexity and data fitne…
Bayesian Generative Active Deep Learning
Toan Tran, Thanh-Toan Do, Ian Reid +1
Deep learning models have demonstrated outstanding performance in several problems, but their training process tends to require immense amounts of computational and human resources…
KL Guided Domain Adaptation
A. Tuan Nguyen, Toan Tran, Yarin Gal +2
Domain adaptation is an important problem and often needed for real-world applications. In this problem, instead of i.i.d. training and testing datapoints, we assume that the sourc…
Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents
Toan Tran, Olivera Kotevska, Li Xiong
Membership inference attacks (MIAs), which enable adversaries to determine whether specific data points were part of a model's training dataset, have emerged as an important framew…
Generalization Bounds for Robust Contrastive Learning: From Theory to Practice
Ngoc N. Tran, Lam Tran, Hoang Phan +5
Contrastive Learning first extracts features from unlabeled data, followed by linear probing with labeled data. Adversarial Contrastive Learning (ACL) integrates Adversarial Traini…
Domain Invariant Representation Learning with Domain Density Transformations
A. Tuan Nguyen, Toan Tran, Yarin Gal +1
Domain generalization refers to the problem where we aim to train a model on data from a set of source domains so that the model can generalize to unseen target domains. Naively tr…
Geo-Llama: Leveraging LLMs for Human Mobility Trajectory Generation with Spatiotemporal Constraints
Siyu Li, Toan Tran, Haowen Lin +5
Generating realistic human mobility data is essential for various application domains, including transportation, urban planning, and epidemic control, as real data is often inacces…
A Bayesian Data Augmentation Approach for Learning Deep Models
Toan Tran, Trung Pham, Gustavo Carneiro +2
Data augmentation is an essential part of the training process applied to deep learning models. The motivation is that a robust training process for deep learning models depends on…
Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning
Anh Tong, Thanh Nguyen-Tang, Dongeun Lee +5
Recent advancements in large language models (LLMs) based on transformer architectures have sparked significant interest in understanding their inner workings. In this paper, we in…
A Theoretically Sound Upper Bound on the Triplet Loss for Improving the Efficiency of Deep Distance Metric Learning
Thanh-Toan Do, Toan Tran, Ian Reid +3
We propose a method that substantially improves the efficiency of deep distance metric learning based on the optimization of the triplet loss function. One epoch of such training p…
Exploiting Domain-Specific Features to Enhance Domain Generalization
Manh-Ha Bui, Toan Tran, Anh Tuan Tran +1
Domain Generalization (DG) aims to train a model, from multiple observed source domains, in order to perform well on unseen target domains. To obtain the generalization capability,…
Improved Training Technique for Shortcut Models
Anh Nguyen, Viet Nguyen, Duc Vu +4
Shortcut models represent a promising, non-adversarial paradigm for generative modeling, uniquely supporting one-step, few-step, and multi-step sampling from a single trained netwo…
SigFormer: Signature Transformers for Deep Hedging
Anh Tong, Thanh Nguyen-Tang, Dongeun Lee +2
Deep hedging is a promising direction in quantitative finance, incorporating models and techniques from deep learning research. While giving excellent hedging strategies, models in…
Improving VANET Simulation Channel Model in an Urban Environment via Calibration Using Real-World Communication Data
Ahmed Gammaa, Seyedmehdi Khaleghian, Toan Tran +1
Wireless communication channels in Vehicular Ad-hoc NETworks (VANETs) suffer from packet losses, which severely influences the performance of their applications. There are several…
Stochastic Multiple Target Sampling Gradient Descent
Hoang Phan, Ngoc Tran, Trung Le +3
Sampling from an unnormalized target distribution is an essential problem with many applications in probabilistic inference. Stein Variational Gradient Descent (SVGD) has been show…
On Learning Domain-Invariant Representations for Transfer Learning with Multiple Sources
Trung Phung, Trung Le, Long Vuong +4
Domain adaptation (DA) benefits from the rigorous theoretical works that study its insightful characteristics and various aspects, e.g., learning domain-invariant representations a…
GeoGNN: Time Series Geo-Localization using Two-Tower Graph Neural Networks
Toan Tran, Waqwoya Abebe, Abhishek Potnis +4
This paper investigates a novel concept of time series geolocalization, where the goal is to infer the geographic origin of each raw time series. Successful geolocalization can pro…
KOPPA: Improving Prompt-based Continual Learning with Key-Query Orthogonal Projection and Prototype-based One-Versus-All
Quyen Tran, Hoang Phan, Lam Tran +4
Drawing inspiration from prompt tuning techniques applied to Large Language Models, recent methods based on pre-trained ViT networks have achieved remarkable results in the field o…
ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation
Ruixuan Liu, Toan Tran, Tianhao Wang +3
As large language models increasingly memorize web-scraped training content, they risk exposing copyrighted or private information. Existing protections require compliance from cra…
TrajGenAgent: A Hierarchical LLM Agent for Human Mobility Trajectory Generation
Siyu Li, Toan Tran, Lingyi Zhao +2
Human mobility data is important for transportation, urban planning, and epidemic control, but large-scale trajectory collection is often costly and privacy-constrained, motivating…
Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models
Duc Anh Nguyen, Tien Ngoc Luu, Tung Pham +1
State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls into states rather than model weights, of…
Reducing Training Time in Cross-Silo Federated Learning using Multigraph Topology
Tuong Do, Binh X. Nguyen, Vuong Pham +4
Federated learning is an active research topic since it enables several participants to jointly train a model without sharing local data. Currently, cross-silo federated learning i…
CASUAL: Conditional Support Alignment for Domain Adaptation with Label Shift
Anh T Nguyen, Lam Tran, Anh Tong +2
Unsupervised domain adaptation (UDA) refers to a domain adaptation framework in which a learning model is trained based on the labeled samples on the source domain and unlabeled on…
Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
Duc Anh Nguyen, Huu Binh Ta, Nhuan Le Duc +2
Sparse Mixture-of-Experts (SMoE) models are scalable and computationally efficient, enabling large increases in model capacity with limited inference overhead. Existing SMoE method…
On Inference Stability for Diffusion Models
Viet Nguyen, Giang Vu, Tung Nguyen Thanh +2
Denoising Probabilistic Models (DPMs) represent an emerging domain of generative models that excel in generating diverse and high-quality images. However, most current training met…
Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training
Toan Tran, Ruixuan Liu, Li Xiong
Large language models (LLMs) have become the backbone of modern natural language processing but pose privacy concerns about leaking sensitive training data. Membership inference at…
Dual-Model Defense: Safeguarding Diffusion Models from Membership Inference Attacks through Disjoint Data Splitting
Bao Q. Tran, Viet Nguyen, Anh Tran +1
Diffusion models have demonstrated remarkable capabilities in image synthesis, but their recently proven vulnerability to Membership Inference Attacks (MIAs) poses a critical priva…
Distributionally Robust Fair Principal Components via Geodesic Descents
Hieu Vu, Toan Tran, Man-Chung Yue +1
Principal component analysis is a simple yet useful dimensionality reduction technique in modern machine learning pipelines. In consequential domains such as college admission, hea…