collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2025

GradientSpace: Unsupervised Data Clustering for Improved Instruction Tuning

Shrihari Sridharan, Deepak Ravikumar, Anand Raghunathan +1

Instruction tuning is one of the key steps required for adapting large language models (LLMs) to a broad spectrum of downstream applications. However, this procedure is difficult b…

cs.LG2025

Coresets from Trajectories: Selecting Data via Correlation of Loss Differences

Manish Nagaraj, Deepak Ravikumar, Kaushik Roy

Deep learning models achieve state-of-the-art performance across domains but face scalability challenges in real-time or resource-constrained scenarios. To address this, we propose…

cs.LG2025

The Easy Path to Robustness: Coreset Selection using Sample Hardness

Pranav Ramesh, Arjun Roy, Deepak Ravikumar +2

Designing adversarially robust models from a data-centric perspective requires understanding which input samples are most crucial for learning resilient features. While coreset sel…

cs.LG2025

Finding the Muses: Identifying Coresets through Loss Trajectories

Manish Nagaraj, Deepak Ravikumar, Efstathia Soufleri +1

Deep learning models achieve state-of-the-art performance across domains but face scalability challenges in real-time or resource-constrained scenarios. To address this, we propose…

cs.LG2025

SAP: Corrective Machine Unlearning with Scaled Activation Projection for Label Noise Robustness

Sangamesh Kodge, Deepak Ravikumar, Gobinda Saha +1

Label corruption, where training samples are mislabeled due to non-expert annotation or adversarial attacks, significantly degrades model performance. Acquiring large, perfectly la…