1 citations · 1 across the 2 of their papers we have counts for
4 papers
Efficiency Adjustments Break the Logarithmic Rank Barrier
Josue Ortega, Geng Zhao, Gabriel Ziegler
We study the expected average rank achieved by the Efficiency-Adjusted Deferred Acceptance (EADA) mechanism in i.i.d.\ matching markets. While student-proposing Deferred Acceptance…
The Distribution of Envy in Matching Markets
Josué Ortega, Gabriel Ziegler, R. Pablo Arribillaga +1
We study the distribution of envy in random matching markets under the Deferred Acceptance (DA) algorithm. Using tools from applied probability, we compute the expected number of p…
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
Julien Piet, Xiao Huang, Dennis Jacob +7
Safety and security remain critical concerns in AI deployment. Despite safety training through reinforcement learning with human feedback (RLHF) [ 32], language models remain vulne…
Toxicity Detection for Free
Zhanhao Hu, Julien Piet, Geng Zhao +2
Current LLMs are generally aligned to follow safety requirements and tend to refuse toxic prompts. However, LLMs can fail to refuse toxic prompts or be overcautious and refuse beni…