activity
20172024
most citedBeyond Performance: Quantifying and Mitigating Label Bias in LLMs

1 citations · 1 across the 2 of their papers we have counts for

collaborators

11 papers

cs.CL20241 cited

Beyond Performance: Quantifying and Mitigating Label Bias in LLMs

Yuval Reif, Roy Schwartz

Large language models (LLMs) have shown remarkable adaptability to diverse tasks, by leveraging context prompts containing instructions, or minimal input-output examples. However,…

cs.CV2022

ARIA: Adversarially Robust Image Attribution for Content Provenance

Maksym Andriushchenko, Xiaoyang Rebecca Li, Geoffrey Oxholm +4

Image attribution -- matching an image back to a trusted source -- is an emerging tool in the fight against online misinformation. Deep visual fingerprinting models have recently b…

cs.LG2020

RobustBench: a standardized adversarial robustness benchmark

Francesco Croce, Maksym Andriushchenko, Vikash Sehwag +5

As a research community, we are still lacking a systematic understanding of the progress on adversarial robustness which often makes it hard to identify the most promising ideas in…

cs.LG2020

Understanding and Improving Fast Adversarial Training

Maksym Andriushchenko, Nicolas Flammarion

A recent line of work focused on making adversarial training computationally efficient for deep learning models. In particular, Wong et al. (2020) showed that -adversa…

cs.LG2020

On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong Baselines

Marius Mosbach, Maksym Andriushchenko, Dietrich Klakow

Fine-tuning pre-trained transformer-based language models such as BERT has become a common practice dominating leaderboards across various NLP benchmarks. Despite the strong empiri…

cs.LG2019

Square Attack: a query-efficient black-box adversarial attack via random search

Maksym Andriushchenko, Francesco Croce, Nicolas Flammarion +1

We propose the Square Attack, a score-based black-box - and -adversarial attack that does not rely on local gradient information and thus is not affected by gradient…