4 papers · 1 filter
Negative Advantages Is a Double-Edged Sword: Calibrating advantages in GRPO for Search Agents
Jiayi Wu, Ruobing Xie, Zeqian Huang +6
Search agents achieve strong question-answering performance through multi-turn interactions with search engines, with Group Relative Policy Optimization (GRPO) being a widely used…
GroupDebate: Enhancing the Efficiency of Multi-Agent Debate Using Group Discussion
Tongxuan Liu, Xingyu Wang, Weizhe Huang +5
In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse NLP tasks. Extensive research has explored how to enhance the logical reasoni…
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
Lei Jiang, Zixun Zhang, Zizhou Wang +4
Large Vision-Language Models (LVLMs) demonstrate exceptional performance across multimodal tasks, yet remain vulnerable to jailbreak attacks that bypass built-in safety mechanisms…
S-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency
Yuting Zeng, Weizhe Huang, Lei Jiang +5
Large language models (LLMs) have demonstrated remarkable capabilities across various natural language processing (NLP) scenarios, but they still face challenges when handling comp…