most citedQuantifying perturbation impacts for large language models

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2025

Visualizing token importance for black-box language models

Paulius Rauba, Qiyao Wei, Mihaela van der Schaar

We consider the problem of auditing black-box large language models (LLMs) to ensure they behave reliably when deployed in production settings, particularly in high-stakes domains…

cs.AI2025

Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?

Matt MacDermott, Qiyao Wei, Rada Djoneva +1

AI systems that output their reasoning in natural language offer an opportunity for safety -- we can \emph{monitor} their chain of thought (CoT) for undesirable reasoning, such as…

cs.AI2025

Semantic-KG: Using Knowledge Graphs to Construct Benchmarks for Measuring Semantic Similarity

Qiyao Wei, Edward Morrell, Lea Goetz +1

Evaluating the open-form textual responses generated by Large Language Models (LLMs) typically requires measuring the semantic similarity of the response to a (human generated) ref…

q-fin.ST2025

Event-Aware Sentiment Factors from LLM-Augmented Financial Tweets: A Transparent Framework for Interpretable Quant Trading

Yueyi Wang, Qiyao Wei

In this study, we wish to showcase the unique utility of large language models (LLMs) in financial semantic annotation and alpha signal discovery. Leveraging a corpus of company-re…

cs.CL2025

Statistical Hypothesis Testing for Auditing Robustness in Language Models

Paulius Rauba, Qiyao Wei, Mihaela van der Schaar

Consider the problem of testing whether the outputs of a large language model (LLM) system change under an arbitrary intervention, such as an input perturbation or changing the mod…

cs.LG20241 cited

Quantifying perturbation impacts for large language models

Paulius Rauba, Qiyao Wei, Mihaela van der Schaar

We consider the problem of quantifying how an input perturbation impacts the outputs of large language models (LLMs), a fundamental task for model reliability and post-hoc interpre…