4 papers
How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning
Jiangwei Chen, Xinyuan Niu, Rachael Hwee Ling Sim +3
Machine unlearning aims to remove the influence of specific forget training data due to privacy, copyright or bias concerns while maintaining the model performance on the remaining…
Incentivizing Time-Aware Fairness in Data Sharing
Jiangwei Chen, Kieu Thao Nguyen Pham, Rachael Hwee Ling Sim +4
In collaborative data sharing and machine learning, multiple parties aggregate their data resources to train a machine learning model with better model performance. However, as the…
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
Gregory Kang Ruey Lau, Xinyuan Niu, Hieu Dao +3
Protecting intellectual property (IP) of text such as articles and code is increasingly important, especially as sophisticated attacks become possible, such as paraphrasing by larg…
Data-Centric AI in the Age of Large Language Models
Xinyi Xu, Zhaoxuan Wu, Rui Qiao +16
This position paper proposes a data-centric viewpoint of AI research, focusing on large language models (LLMs). We start by making the key observation that data is instrumental in…