8 papers
A Sober Look at Agentic Misalignment in Automated Workflows
Wenqian Ye, Bo Yuan, Zhichao Xu +4
We study a class of emergent misalignment in multi-agent systems (MAS), with a focus on automated workflows, which we refer to agentic misalignment. Although these systems can solv…
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
Wenqian Ye, Bohan Liu, Guangtao Zheng +6
Spurious bias, a tendency to exploit spurious correlations between superficial input attributes and prediction targets, has revealed a severe robustness pitfall in classical machin…
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
Wenqian Ye, Di Wang, Guangtao Zheng +2
Large vision-language models, such as CLIP, have shown strong zero-shot classification performance by aligning images and text in a shared embedding space. However, CLIP models oft…
Rectifying Shortcut Behaviors in Preference-based Reward Learning
Wenqian Ye, Guangtao Zheng, Aidong Zhang
In reinforcement learning from human feedback, preference-based reward models play a central role in aligning large language models to human-aligned behavior. However, recent studi…
The Clever Hans Mirage: A Comprehensive Survey on Spurious Correlations in Machine Learning
Wenqian Ye, Luyang Jiang, Eric Xie +16
Back in the early 20th century, a horse named Hans appeared to perform arithmetic and other intellectual tasks during exhibitions in Germany, while it actually relied solely on inv…
Improving Group Robustness on Spurious Correlation via Evidential Alignment
Wenqian Ye, Guangtao Zheng, Aidong Zhang
Deep neural networks often learn and rely on spurious correlations, i.e., superficial associations between non-causal features and the targets. For instance, an image classifier ma…