activity
20232026
most citedWaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models

1 citations · 2 across the 8 of their papers we have counts for

collaborators

19 papers

cs.CY2026

Uncertainty-based Debiasing and Unlearning for Decontamination

Guangzhi Sun, Xiao Zhan, Mark Gales

Benchmark-based evaluation is the dominant paradigm for assessing large language model (LLM) capabilities, yet data contamination inflates reported performance and undermines fair…

cs.CL2024

SkillAggregation: Reference-free LLM-Dependent Aggregation

Guangzhi Sun, Anmol Kagrecha, Potsawee Manakul +2

Large Language Models (LLMs) are increasingly used to assess NLP tasks due to their ability to generate human-like judgments. Single LLMs were used initially, however, recent work…

cs.CL2024

Finetuning LLMs for Comparative Assessment Tasks

Vatsal Raina, Adian Liusie, Mark Gales

Automated assessment in natural language generation is a challenging task. Instruction-tuned large language models (LLMs) have shown promise in reference-free evaluation, particula…

cs.SD2024

Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models

Vyas Raina, Mark Gales

Speech enabled foundation models, either in the form of flexible speech recognition based systems or audio-prompted large language models (LLMs), are becoming increasingly popular.…

cs.CL2024

Cross-Lingual Transfer Learning for Speech Translation

Rao Ma, Mengjie Qian, Yassir Fathullah +3

There has been increasing interest in building multilingual foundation models for NLP and speech research. This paper examines how to expand the speech translation capability of th…

cs.CL2024

CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models

Guangzhi Sun, Potsawee Manakul, Adian Liusie +4

Multimodal foundation models are prone to hallucination, generating outputs that either contradict the input or are not grounded by factual information. Given the diversity in arch…