1 citations · 1 across the 13 of their papers we have counts for
3 papers · 1 filter
Understanding Scam Trends and Rail Paths from Reddit Self-Disclosure Narratives
Yangjun Zhang, Mirko Bottarelli, Mark Hooper +1
Online scam behavior is inherently multi-stage, and the lifecycle includes temporally ordered rails and events rather than isolated signals. Existing works analyze characteristics…
Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges
Riya Tapwal, Abhishek Kumar, Carsten Maple
Large language models (LLMs) are increasingly used as automatic judges for summarization and dialogue evaluation. Prior work has documented biases such as position, verbosity, and…
Representation Noising: A Defence Mechanism Against Harmful Finetuning
Domenic Rosati, Jan Wehner, Kai Williams +7
Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release o…