10 papers · 1 filter
The Chicken and Egg Dilemma: Co-optimizing Data and Model Configurations for LLMs
Zhiliang Chen, Alfred Wei Lun Leong, Shao Yong Ong +6
Co-optimizing data and model configurations for training LLMs presents a classic chicken-and-egg dilemma: The best training data configuration (e.g., data mixture) for a downstream…
Incentivizing Time-Aware Fairness in Data Sharing
Jiangwei Chen, Kieu Thao Nguyen Pham, Rachael Hwee Ling Sim +4
In collaborative data sharing and machine learning, multiple parties aggregate their data resources to train a machine learning model with better model performance. However, as the…
WaterDrum: Watermarking for Data-centric Unlearning Metric
Xinyang Lu, Xinyuan Niu, Gregory Kang Ruey Lau +7
Large language model (LLM) unlearning is critical in real-world applications where it is necessary to efficiently remove the influence of private, copyrighted, or harmful data from…
Confidence Elicitation: A New Attack Vector for Large Language Models
Brian Formento, Chuan Sheng Foo, See-Kiong Ng
A fundamental issue in deep learning has been adversarial robustness. As these systems have scaled, such issues have persisted. Currently, large language models (LLMs) with billion…
DUET: Optimizing Training Data Mixtures via Feedback from Unseen Evaluation Tasks
Zhiliang Chen, Gregory Kang Ruey Lau, Chuan-Sheng Foo +1
The performance of an LLM depends heavily on the relevance of its training data to the downstream evaluation task. However, in practice, the data involved in an unseen evaluation t…
Data Distribution Valuation
Xinyi Xu, Shuaiqi Wang, Chuan-Sheng Foo +2
Data valuation is a class of techniques for quantitatively assessing the value of data for applications like pricing in data marketplaces. Existing data valuation methods define a…