12 papers
APPO: Agentic Procedural Policy Optimization
Xucong Wang, Ziyu Ma, Yong Wang +5
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing metho…
Semantically Structured Mixture-of-Experts for Compositional Robotic Manipulation
Chengyu Deng, Guanqi Chen, Yizhou Chen +4
Diffusion-based policies have established a new standard for precise robotic manipulation but face a critical scalability bottleneck: high-performance models are computationally ex…
CaptchaMind: Training CAPTCHA Solvers via Reinforcement Learning with Explicit Reasoning Supervision
Pengcheng Wang, Haoxiang Liu, Yang Dai +5
CAPTCHAs are widely deployed as human verification mechanisms and frequently block intelligent agents from completing end-to-end automation in real-world web environments. Solving…
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
Lecheng Yan, Ruizhe Li, Xicheng Han +5
Tool-using LLM agents increasingly rely on external tools to make consequential decisions, yet most existing agent-security benchmarks and defenses implicitly assume that tool feed…
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
Peiqin Lin, Chenyang Lyu, Wenjiang Luo +22
Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities. However, these benchmarks prio…
Anchored Supervised Fine-Tuning
He Zhu, Junyou Su, Peng Lai +4
Post-training of large language models involves a fundamental trade-off between supervised fine-tuning (SFT), which efficiently mimics demonstrations but tends to memorize, and rei…