1 citations · 1 across the 1 of their papers we have counts for
4 papers
Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models
Schrasing Tong, Eliott Zemour, Jessica Lu +2
Although large language models (LLMs) have demonstrated their effectiveness in a wide range of applications, they have also been observed to perpetuate unwanted biases present in t…
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
Francisco Eiras, Eliott Zemour, Eric Lin +1
Large Language Model (LLM) based judges form the underpinnings of key safety evaluation processes such as offline benchmarking, automated red-teaming, and online guardrailing. This…
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
Blazej Manczak, Eliott Zemour, Eric Lin +1
Deploying language models (LMs) necessitates outputs to be both high-quality and compliant with safety guidelines. Although Inference-Time Guardrails (ITG) offer solutions that shi…
Does fine-tuning GPT-3 with the OpenAI API leak personally-identifiable information?
Albert Yu Sun, Eliott Zemour, Arushi Saxena +4
Machine learning practitioners often fine-tune generative pre-trained models like GPT-3 to improve model performance at specific tasks. Previous works, however, suggest that fine-t…