5 papers
WMAttack: Automated Attack Search for Adversarial Evaluation of World-Model Agents
Zhixiang Guo, Siyuan Liang, Shi Fu +4
Despite the growing use of world models as decision-making agents, their adversarial robustness remains underexplored due to the lack of dedicated automated evaluation methods. A k…
ATAC: Augmentation-Based Test-Time Adversarial Correction for CLIP
Linxiang Su, András Balogh
Despite its remarkable success in zero-shot image-text matching, CLIP remains highly vulnerable to adversarial perturbations on images. As adversarial fine-tuning is prohibitively…
When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World Models
Zhixiang Guo, Siyuan Liang, Andras Balogh +4
Generative world models (WMs) are increasingly used to synthesize controllable, sensor-conditioned driving videos, yet their reliance on physical priors exposes novel attack surfac…
Verification of the Implicit World Model in a Generative Model via Adversarial Sequences
András Balogh, Márk Jelasity
Generative sequence models are typically trained on sample sequences from natural or formal languages. It is a crucial question whether -- or to what extent -- sample-based trainin…
How not to Stitch Representations to Measure Similarity: Task Loss Matching versus Direct Matching
András Balogh, Márk Jelasity
Measuring the similarity of the internal representations of deep neural networks is an important and challenging problem. Model stitching has been proposed as a possible approach,…