10 papers
Towards Irreversible Machine Unlearning for Diffusion Models
Xun Yuan, Zilong Zhao, Jiayu Li +3
Diffusion models are renowned for their state-of-the-art performance in generating synthetic images. However, concerns related to safety, privacy, and copyright highlight the need…
Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day
Milad Abdollahzadeh, Abdul Raheem, Zilong Zhao +5
Tabular instruction tuning has emerged as a promising research direction for improving LLMs understanding of tabular data. However, the majority of existing works only consider que…
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought
Tencent Hunyuan Team, Ao Liu, Botong Zhou +248
As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mam…
Generating Synthetic Data with Formal Privacy Guarantees: State of the Art and the Road Ahead
Viktor Schlegel, Anil A Bharath, Zilong Zhao +1
Privacy-preserving synthetic data offers a promising solution to harness segregated data in high-stakes domains where information is compartmentalized for regulatory, privacy, or i…
TabTreeFormer: Tabular Data Generation Using Hybrid Tree-Transformer
Jiayu Li, Bingyin Zhao, Zilong Zhao +3
Transformers have shown impressive results in tabular data generation. However, they lack domain-specific inductive biases which are critical for preserving the intrinsic character…
TAEGAN: Generating Synthetic Tabular Data For Data Augmentation
Jiayu Li, Zilong Zhao, Kevin Yee +2
Synthetic tabular data generation has gained significant attention for its potential in data augmentation and privacy-preserving data sharing. While recent methods like diffusion a…