3 citations · 7 across the 5 of their papers we have counts for
6 papers
Hubble: a Model Suite to Advance the Study of LLM Memorization
Johnny Tian-Zheng Wei, Ameya Godbole, Mohammad Aflah Khan +7
We present Hubble, a suite of fully open-source large language models (LLMs) for the scientific study of LLM memorization. Hubble models come in standard and perturbed variants: st…
Interrogating LLM design under a fair learning doctrine
Johnny Tian-Zheng Wei, Maggie Wang, Ameya Godbole +2
The current discourse on large language models (LLMs) and copyright largely takes a "behavioral" perspective, focusing on model outputs and evaluating whether they are substantiall…
Proving membership in LLM pretraining data via data watermarks
Johnny Tian-Zheng Wei, Ryan Yixiang Wang, Robin Jia
Detecting whether copyright holders' works were used in LLM pretraining is poised to be an important problem. This work proposes using data watermarks to enable principled detectio…
Operationalizing content moderation "accuracy" in the Digital Services Act
Johnny Tian-Zheng Wei, Frederike Zufall, Robin Jia
The Digital Services Act, recently adopted by the EU, requires social media platforms to report the "accuracy" of their automated content moderation systems. The colloquial term is…
Searching for a higher power in the human evaluation of MT
Johnny Tian-Zheng Wei, Tom Kocmi, Christian Federmann
In MT evaluation, pairwise comparisons are conducted to identify the better system. In conducting the comparison, the experimenter must allocate a budget to collect Direct Assessme…
The statistical advantage of automatic NLG metrics at the system level
Johnny Tian-Zheng Wei, Robin Jia
Estimating the expected output quality of generation systems is central to NLG. This paper qualifies the notion that automatic metrics are not as good as humans in estimating syste…