17 citations · 40 across the 23 of their papers we have counts for
23 papers
Data value estimation on private gradients
Zijian Zhou, Xinyi Xu, Daniela Rus +1
For gradient-based machine learning (ML) methods commonly adopted in practice such as stochastic gradient descent, the de facto differential privacy (DP) technique is perturbing th…
Self-Interested Agents in Collaborative Machine Learning: An Incentivized Adaptive Data-Centric Framework
Nithia Vijayan, Bryan Kian Hsiang Low
We propose a framework for adaptive data-centric collaborative machine learning among self-interested agents, coordinated by an arbiter. Designed to handle the incremental nature o…
Global-to-Local Support Spectrums for Language Model Explainability
Lucas Agussurja, Xinyang Lu, Bryan Kian Hsiang Low
Existing sample-based methods, like influence functions and representer points, measure the importance of a training point by approximating the effect of its removal from training.…
TRACE: TRansformer-based Attribution using Contrastive Embeddings in LLMs
Cheng Wang, Xinyang Lu, See-Kiong Ng +1
The rapid evolution of large language models (LLMs) represents a substantial leap forward in natural language understanding and generation. However, alongside these advancements co…
Data-Centric AI in the Age of Large Language Models
Xinyi Xu, Zhaoxuan Wu, Rui Qiao +16
This position paper proposes a data-centric viewpoint of AI research, focusing on large language models (LLMs). We start by making the key observation that data is instrumental in…
Helpful or Harmful Data? Fine-tuning-free Shapley Attribution for Explaining Language Model Predictions
Jingtan Wang, Xiaoqiang Lin, Rui Qiao +2
The increasing complexity of foundational models underscores the necessity for explainability, particularly for fine-tuning, the most widely used training method for adapting model…