Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Position: Multi-Agent Systems Should Prioritize Concurrency Control
Xin Yang, Letian Li, Zimo Ji +2
LLM-based multi-agent systems (MAS) promise scalable collaboration, yet adding agents often reduces reliability. This position paper argues that many MAS failures are fundamentally…
cs.AI2025
INTEGRALBENCH: Benchmarking LLMs with Definite Integral Problems
Bintao Tang, Xin Yang, Yuhao Wang +3
We present INTEGRALBENCH, a focused benchmark designed to evaluate Large Language Model (LLM) performance on definite integral problems. INTEGRALBENCH provides both symbolic and nu…
cs.AI2025
Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges
Zimo Ji, Daoyuan Wu, Wenyuan Jiang +3
Capture-the-Flag (CTF) competitions are crucial for cybersecurity education and training. As large language models (LLMs) evolve, there is increasing interest in their ability to a…