3 papers
cs.AI2026
History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions
Alberto G. RodrÃguez Salgado
Frontier LLMs are increasingly deployed as agents that pick the next action after a long log of prior tool calls produced by the same or a different model. We ask a simple safety q…
cs.LG2026
From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning
Alberto G. Rodriguez Salgado
How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce \textsc{MazeBench}, a benchmark of 110 p…
cs.CV2024
Global Average Feature Augmentation for Robust Semantic Segmentation with Transformers
Alberto Gonzalo Rodriguez Salgado, Maying Shen, Philipp Harzig +2
Robustness to out-of-distribution data is crucial for deploying modern neural networks. Recently, Vision Transformers, such as SegFormer for semantic segmentation, have shown impre…