4 citations · 9 across the 10 of their papers we have counts for
10 papers
Geometry-Aware Rotary Position Embedding for Consistent Video World Model
Chendong Xiang, Jiajun Liu, Jintao Zhang +7
Predictive world models that simulate future observations under explicit camera control are fundamental to interactive AI. Despite rapid advances, current systems lack spatial pers…
Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents
Renhua Ding, Xiao Yang, Zhengwei Fang +3
Large Vision-Language Models (LVLMs) empower autonomous mobile agents, yet their security under realistic mobile deployment constraints remains underexplored. While agents are vuln…
Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
Yichi Zhang, Yao Huang, Yifan Wang +10
The trustworthiness of Multimodal Large Language Models (MLLMs) remains an intense concern despite the significant progress in their capabilities. Existing evaluation and mitigatio…
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
Xiao Yang, Jiawei Chen, Jun Luo +4
The emergence of multimodal LLM-based agents (MLAs) has transformed interaction paradigms by seamlessly integrating vision, language, action and dynamic environments, enabling unpr…
Exploring the Secondary Risks of Large Language Models
Jiawei Chen, Zhengwei Fang, Yu Tian +4
Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior…
STAIR: Improving Safety Alignment with Introspective Reasoning
Yichi Zhang, Siyuan Zhang, Yao Huang +7
Ensuring the safety and harmlessness of Large Language Models (LLMs) has become equally critical as their performance in applications. However, existing safety alignment methods ty…