papers

Publications (31)

stat.AP2026

Multilevel Regression and Poststratification Interface: An Application to Track Community-level COVID-19 Viral Transmission

Yajuan Si, Toan Tran, Jonah Gabry +2

We present a novel Bayesian workflow for multilevel regression and poststratification (MRP), introducing extensions to time-varying data and granular geography and publicly availab…

cs.CV2025

Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts

Viet Nguyen, Anh Nguyen, Trung Dao +4

The escalating demand for real-time image synthesis has driven significant advancements in one-step diffusion models, which inherently offer expedited generation speeds compared to…

cs.LG2026

Differentially Private Synthetic Data via APIs 4: Tabular Data

Toan Tran, Arturs Backurs, Zinan Lin +3

This paper investigates the problem of generating synthetic tabular data with differential privacy (DP) guarantees, enabling data sharing in sensitive domains. Despite extensive st…

cs.LG2021

Learning Compositional Sparse Gaussian Processes with a Shrinkage Prior

Anh Tong, Toan Tran, Hung Bui +1

Choosing a proper set of kernel functions is an important problem in learning Gaussian Process (GP) models since each kernel structure has different model complexity and data fitne…

cs.LG2019

Bayesian Generative Active Deep Learning

Toan Tran, Thanh-Toan Do, Ian Reid +1

Deep learning models have demonstrated outstanding performance in several problems, but their training process tends to require immense amounts of computational and human resources…

cs.LG2022

KL Guided Domain Adaptation

A. Tuan Nguyen, Toan Tran, Yarin Gal +2

Domain adaptation is an important problem and often needed for real-world applications. In this problem, instead of i.i.d. training and testing datapoints, we assume that the sourc…

cs.CR2026

Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents

Toan Tran, Olivera Kotevska, Li Xiong

Membership inference attacks (MIAs), which enable adversaries to determine whether specific data points were part of a model's training dataset, have emerged as an important framew…

cs.LG2025

Generalization Bounds for Robust Contrastive Learning: From Theory to Practice

Ngoc N. Tran, Lam Tran, Hoang Phan +5

Contrastive Learning first extracts features from unlabeled data, followed by linear probing with labeled data. Adversarial Contrastive Learning (ACL) integrates Adversarial Traini…

cs.LG2022

Domain Invariant Representation Learning with Domain Density Transformations

A. Tuan Nguyen, Toan Tran, Yarin Gal +1

Domain generalization refers to the problem where we aim to train a model on data from a set of source domains so that the model can generalize to unseen target domains. Naively tr…

cs.AI2025

Geo-Llama: Leveraging LLMs for Human Mobility Trajectory Generation with Spatiotemporal Constraints

Siyu Li, Toan Tran, Haowen Lin +5

Generating realistic human mobility data is essential for various application domains, including transportation, urban planning, and epidemic control, as real data is often inacces…

cs.CV2017

A Bayesian Data Augmentation Approach for Learning Deep Models

Toan Tran, Trung Pham, Gustavo Carneiro +2

Data augmentation is an essential part of the training process applied to deep learning models. The motivation is that a robust training process for deep learning models depends on…

cs.LG2025

Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning

Anh Tong, Thanh Nguyen-Tang, Dongeun Lee +5

Recent advancements in large language models (LLMs) based on transformer architectures have sparked significant interest in understanding their inner workings. In this paper, we in…

cs.CV2019

A Theoretically Sound Upper Bound on the Triplet Loss for Improving the Efficiency of Deep Distance Metric Learning

Thanh-Toan Do, Toan Tran, Ian Reid +3

We propose a method that substantially improves the efficiency of deep distance metric learning based on the optimization of the triplet loss function. One epoch of such training p…

cs.LG2021

Exploiting Domain-Specific Features to Enhance Domain Generalization

Manh-Ha Bui, Toan Tran, Anh Tuan Tran +1

Domain Generalization (DG) aims to train a model, from multiple observed source domains, in order to perform well on unseen target domains. To obtain the generalization capability,…

cs.CV2025

Improved Training Technique for Shortcut Models

Anh Nguyen, Viet Nguyen, Duc Vu +4

Shortcut models represent a promising, non-adversarial paradigm for generative modeling, uniquely supporting one-step, few-step, and multi-step sampling from a single trained netwo…

cs.LG2023

SigFormer: Signature Transformers for Deep Hedging

Anh Tong, Thanh Nguyen-Tang, Dongeun Lee +2

Deep hedging is a promising direction in quantitative finance, incorporating models and techniques from deep learning research. While giving excellent hedging strategies, models in…

cs.ET2025

Improving VANET Simulation Channel Model in an Urban Environment via Calibration Using Real-World Communication Data

Ahmed Gammaa, Seyedmehdi Khaleghian, Toan Tran +1

Wireless communication channels in Vehicular Ad-hoc NETworks (VANETs) suffer from packet losses, which severely influences the performance of their applications. There are several…

cs.LG2023

Stochastic Multiple Target Sampling Gradient Descent

Hoang Phan, Ngoc Tran, Trung Le +3

Sampling from an unnormalized target distribution is an essential problem with many applications in probabilistic inference. Stein Variational Gradient Descent (SVGD) has been show…

cs.LG2021

On Learning Domain-Invariant Representations for Transfer Learning with Multiple Sources

Trung Phung, Trung Le, Long Vuong +4

Domain adaptation (DA) benefits from the rigorous theoretical works that study its insightful characteristics and various aspects, e.g., learning domain-invariant representations a…

cs.LG2026

GeoGNN: Time Series Geo-Localization using Two-Tower Graph Neural Networks

Toan Tran, Waqwoya Abebe, Abhishek Potnis +4

This paper investigates a novel concept of time series geolocalization, where the goal is to infer the geographic origin of each raw time series. Successful geolocalization can pro…

cs.LG2024

KOPPA: Improving Prompt-based Continual Learning with Key-Query Orthogonal Projection and Prototype-based One-Versus-All

Quyen Tran, Hoang Phan, Lam Tran +4

Drawing inspiration from prompt tuning techniques applied to Large Language Models, recent methods based on pre-trained ViT networks have achieved remarkable results in the field o…

cs.CR2025

ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation

Ruixuan Liu, Toan Tran, Tianhao Wang +3

As large language models increasingly memorize web-scraped training content, they risk exposing copyrighted or private information. Existing protections require compliance from cra…

cs.AI2026

TrajGenAgent: A Hierarchical LLM Agent for Human Mobility Trajectory Generation

Siyu Li, Toan Tran, Lingyi Zhao +2

Human mobility data is important for transportation, urban planning, and epidemic control, but large-scale trajectory collection is often costly and privacy-constrained, motivating…

cs.LG2026

Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models

Duc Anh Nguyen, Tien Ngoc Luu, Tung Pham +1

State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls into states rather than model weights, of…

cs.LG2023

Reducing Training Time in Cross-Silo Federated Learning using Multigraph Topology

Tuong Do, Binh X. Nguyen, Vuong Pham +4

Federated learning is an active research topic since it enables several participants to jointly train a model without sharing local data. Currently, cross-silo federated learning i…

cs.LG2024

CASUAL: Conditional Support Alignment for Domain Adaptation with Label Shift

Anh T Nguyen, Lam Tran, Anh Tong +2

Unsupervised domain adaptation (UDA) refers to a domain adaptation framework in which a learning model is trained based on the labeled samples on the source domain and unlabeled on…

cs.LG2026

Selective Sinkhorn Routing for Improved Sparse Mixture of Experts

Duc Anh Nguyen, Huu Binh Ta, Nhuan Le Duc +2

Sparse Mixture-of-Experts (SMoE) models are scalable and computationally efficient, enabling large increases in model capacity with limited inference overhead. Existing SMoE method…

cs.CV2024

On Inference Stability for Diffusion Models

Viet Nguyen, Giang Vu, Tung Nguyen Thanh +2

Denoising Probabilistic Models (DPMs) represent an emerging domain of generative models that excel in generating diverse and high-quality images. However, most current training met…

cs.LG2025

Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training

Toan Tran, Ruixuan Liu, Li Xiong

Large language models (LLMs) have become the backbone of modern natural language processing but pose privacy concerns about leaking sensitive training data. Membership inference at…

cs.LG2024

Dual-Model Defense: Safeguarding Diffusion Models from Membership Inference Attacks through Disjoint Data Splitting

Bao Q. Tran, Viet Nguyen, Anh Tran +1

Diffusion models have demonstrated remarkable capabilities in image synthesis, but their recently proven vulnerability to Membership Inference Attacks (MIAs) poses a critical priva…

cs.LG2022

Distributionally Robust Fair Principal Components via Geodesic Descents

Hieu Vu, Toan Tran, Man-Chung Yue +1

Principal component analysis is a simple yet useful dimensionality reduction technique in modern machine learning pipelines. In consequential domains such as college admission, hea…