7 papers
AcquisitionSynthesis: Targeted Data Generation using Acquisition Functions
Ishika Agarwal, Sofia Stoica, Emre Can Acikgoz +4
Data quality remains a critical bottleneck in developing capable, competitive models. Researchers have explored many ways to generate top quality samples. Some works rely on reject…
Language Specific Knowledge: Do Models Know Better in X than in English?
Ishika Agarwal, Nimet Beyza Bozdag, Nisval Patel +1
Often, multilingual language models are trained with the objective to map semantically similar content (in different languages) in the same latent space. In this paper, we show a n…
A Rising Tide Lifts All Boats: MTQE Rewards for Idioms Improve General Translation Quality
Ishika Agarwal, Zhenlin He, Dhruva Patil +1
Non-compositional expressions (e.g., idioms, proverbs, and metaphors) pose significant challenges for neural machine translation systems because their meanings cannot be derived fr…
Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning Data
Ishika Agarwal, Dilek Hakkani-Tür
Influence functions provide crucial insights into model training, but existing methods suffer from large computational costs and limited generalization. Particularly, recent works…
Tree-of-Debate: Multi-Persona Debate Trees Elicit Critical Thinking for Scientific Comparative Analysis
Priyanka Kargupta, Ishika Agarwal, Tal August +1
With the exponential growth of research facilitated by modern technology and improved accessibility, scientific discoveries have become increasingly fragmented within and across fi…
DELIFT: Data Efficient Language model Instruction Fine Tuning
Ishika Agarwal, Krishnateja Killamsetty, Lucian Popa +1
Fine-tuning large language models (LLMs) is essential for enhancing their performance on specific tasks but is often resource-intensive due to redundant or uninformative data. To a…