activity
20242026
collaborators

7 papers

cs.AI2026

Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem

Weixun Wang, XiaoXiao Xu, Wanhe An +86

Agentic crafting requires LLMs to operate in real-world environments over multiple turns by taking actions, observing outcomes, and iteratively refining artifacts. Despite its impo…

cs.CV2026

CMOOD: Concept-based Multi-label OOD Detection

Zhendong Liu, Yi Nian, Yuehan Qin +4

How can models effectively detect out-of-distribution (OOD) samples in complex, multi-label settings without extensive retraining? Existing OOD detection methods struggle to captur…

cs.CV2025

Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based Attacks

Jiawei Wang, Yushen Zuo, Yuanjun Chai +4

Vision-Language Models (VLMs) extend the capabilities of Large Language Models (LLMs) by incorporating visual information, yet they remain vulnerable to jailbreak attacks, especial…

cs.CR2025

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models

Baolin Zheng, Guanlin Chen, Hongqiong Zhong +12

Despite their remarkable achievements and widespread adoption, Multimodal Large Language Models (MLLMs) have revealed significant security vulnerabilities, highlighting the urgent…

cs.AI2025

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

Baihui Zheng, Boren Zheng, Kerui Cao +9

Despite the remarkable proficiency of \textit{Large Reasoning Models} (LRMs) in handling complex reasoning tasks, their reliability in safety-critical scenarios remains uncertain.…

cs.CV2025

PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment

Zhendong Liu, Yuanbi Nie, Yingshui Tan +6

Benefiting from the powerful capabilities of Large Language Models (LLMs), pre-trained visual encoder models connected to LLMs form Vision Language Models (VLMs). However, recent r…