4 papers
RoPoLL: Robust Panel of LLM Judges
Anish Acharya, Kris W Pan, Brian Verkhovsky
The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its statistical behavior remains p…
Understanding Contrastive Representation Learning from Positive Unlabeled (PU) Data
Anish Acharya, Li Jing, Bhargav Bhushanam +4
Pretext Invariant Representation Learning (PIRL) followed by Supervised Fine-Tuning (SFT) has become a standard paradigm for learning with limited labels. We extend this approach t…
Geometric Median Matching for Robust k-Subset Selection from Noisy Data
Anish Acharya, Sujay Sanghavi, Alexandros G. Dimakis +1
Data pruning -- the combinatorial task of selecting a small and representative subset from a large dataset, is crucial for mitigating the enormous computational costs associated wi…
Geometric Median (GM) Matching for Robust Data Pruning
Anish Acharya, Inderjit S Dhillon, Sujay Sanghavi
Large-scale data collections in the wild, are invariably noisy. Thus developing data pruning strategies that remain robust even in the presence of corruption is critical in practic…