11 papers
Seq2Synth: Benchmarking Temporal Fidelity in Synthetic Sequential Tabular Data
Kiwan Kwon, Kangmin Kim, Hojin Lee +5
Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and data-driven research, but evaluating their fidelity remains difficult because tempor…
Adaptive and Robust Watermark for Generative Tabular Data
Dung Daniel Ngo, Archan Ray, Akshay Seshadri +6
In recent years, watermarking generative tabular data has become a prominent framework to protect against the misuse of synthetic data. However, while most prior work in watermarki…
Dynamic Linear Coregionalization for Realistic Synthetic Multivariate Time Series
Annita Vapsi, Penghang Liu, Saheed Obitayo +8
Synthetic data is essential for training foundation models for time series (FMTS), but most generators assume static correlations, and are typically missing realistic inter-channel…
TS-Agent: Understanding and Reasoning Over Raw Time Series via Iterative Insight Gathering
Penghang Liu, Elizabeth Fons, Annita Vapsi +5
Large language models (LLMs) exhibit strong symbolic and compositional reasoning, yet they struggle with time series question answering as the data is typically transformed into an…
Explicit Group Sparse Projection with Applications to Deep Learning and NMF
Riyasat Ohib, Nicolas Gillis, Niccolò Dalmasso +3
We design a new sparse projection method for a set of vectors that guarantees a desired average sparsity level measured leveraging the popular Hoyer measure (an affine function of…
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
Rongzhe Wei, Peizhi Niu, Hans Hao-Hsun Hsu +9
Machine unlearning techniques aim to mitigate unintended memorization in large language models (LLMs). However, existing approaches predominantly focus on the explicit removal of i…