3 papers
cs.LG2025
Dependency-aware synthetic tabular data generation
Chaithra Umesh, Kristian Schultz, Manjunath Mahendra +2
Synthetic tabular data is increasingly used in privacy-sensitive domains such as health care, but existing generative models often fail to preserve inter-attribute relationships. I…
cs.LG2025
Handling Missing Data in Downstream Tasks With Distribution-Preserving Guarantees
Rahul Bordoloi, Clémence Réda, Saptarshi Bej +1
Missing feature values are a significant hurdle for downstream machine-learning tasks such as classification. However, imputation methods for classification might be time-consuming…
cs.LG2024
Preserving logical and functional dependencies in synthetic tabular data
Chaithra Umesh, Kristian Schultz, Manjunath Mahendra +2
Dependencies among attributes are a common aspect of tabular data. However, whether existing tabular data generation algorithms preserve these dependencies while generating synthet…