5 papers
Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
Atoosa Chegini, Hamid Kazemi, Garrett Souza +5
Reasoning has become a central paradigm for large language models (LLMs), consistently boosting accuracy across diverse benchmarks. Yet its suitability for precision-sensitive task…
RePanda: Pandas-powered Tabular Verification and Reasoning
Atoosa Malemir Chegini, Keivan Rezaei, Hamid Eghbalzadeh +1
Fact-checking tabular data is essential for ensuring the accuracy of structured information. However, existing methods often rely on black-box models with opaque reasoning. We intr…
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
Atoosa Chegini, Hamid Kazemi, Iman Mirzadeh +5
In Large Language Model (LLM) development, Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning models with human values and preferences. RLHF traditionally re…
Data-Centric Debugging: mitigating model failures via targeted data collection
Sahil Singla, Atoosa Malemir Chegini, Mazda Moayeri +1
Deep neural networks can be unreliable in the real world when the training set does not adequately cover all the settings where they are deployed. Focusing on image classification,…
InForecaster: Forecasting Influenza Hemagglutinin Mutations Through the Lens of Anomaly Detection
Ali Garjani, Atoosa Malemir Chegini, Mohammadreza Salehi +6
The influenza virus hemagglutinin is an important part of the virus attachment to the host cells. The hemagglutinin proteins are one of the genetic regions of the virus with a high…