activity
20242026
collaborators

6 papers

cs.AI2026

Harnessing non-adversarial robustness in large language models

Qinghua Zhou, Ellina Aleshina, Andrey Lovyagin +6

The work presents an approach for addressing the challenge of robustness in Large Language Models (LLMs) to alterations and potential errors caused by semantically similar but text…

cs.CV2025

Staining and locking computer vision models without retraining

Oliver J. Sutton, Qinghua Zhou, George Leete +2

We introduce new methods of staining and locking computer vision models, to protect their owners' intellectual property. Staining, also known as watermarking, embeds secret behavio…

cs.LO2025

StepProof: Step-by-step verification of natural language mathematical proofs

Xiaolin Hu, Qinghua Zhou, Bogdan Grechuk +1

Interactive theorem provers (ITPs) are powerful tools for the formal verification of mathematical proofs down to the axiom level. However, their lack of a natural language interfac…

cs.LG2024

The Boundaries of Verifiable Accuracy, Robustness, and Generalisation in Deep Learning

Alexander Bastounis, Alexander N. Gorban, Anders C. Hansen +5

In this work, we assess the theoretical limitations of determining guaranteed stability and accuracy of neural networks in classification tasks. We consider classical distribution-…

cs.AI2024

Stealth edits to large language models

Oliver J. Sutton, Qinghua Zhou, Wei Wang +4

We reveal the theoretical foundations of techniques for editing large language models, and present new methods which can do so without requiring retraining. Our theoretical insight…

cs.LG2024

How adversarial attacks can disrupt seemingly stable accurate classifiers

Oliver J. Sutton, Qinghua Zhou, Ivan Y. Tyukin +3

Adversarial attacks dramatically change the output of an otherwise accurate learning system using a seemingly inconsequential modification to a piece of input data. Paradoxically,…