3 papers
cs.MA2026
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
Yuzhe Zhang, Feiran Liu, Yi Shan +8
Large language models are increasingly deployed in multi-agent systems to overcome context limitations by distributing information across agents. Yet whether agents can reliably co…
cs.AI2025
INTEGRALBENCH: Benchmarking LLMs with Definite Integral Problems
Bintao Tang, Xin Yang, Yuhao Wang +3
We present INTEGRALBENCH, a focused benchmark designed to evaluate Large Language Model (LLM) performance on definite integral problems. INTEGRALBENCH provides both symbolic and nu…
cs.CR2025
Towards Provable (In)Secure Model Weight Release Schemes
Xin Yang, Bintao Tang, Yuhao Wang +3
Recent secure weight release schemes claim to enable open-source model distribution while protecting model ownership and preventing misuse. However, these approaches lack rigorous…