2 papers
cs.AI2026
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
Ruipeng Wang, Yuxin Chen, Yukai Wang +9
Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployment…
cs.LG2025
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
Junfeng Fang, Yukai Wang, Ruipeng Wang +5
The rapid advancement of multi-modal large reasoning models (MLRMs) -- enhanced versions of multimodal language models (MLLMs) equipped with reasoning capabilities -- has revolutio…