3 papers
cs.SE2026
Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents
Alibek Kaliyev, Artem Maryanskyy
Agents that synthesize their own tools ship a second artifact alongside each answer: a software library that future tasks reuse, compose, and depend on. Task completion (TC) certif…
cs.MA2026
How Much Coordination Gain Is Real? A Paired Noise-Floor Protocol for Multi-Agent LLM Benchmarks
Alibek T Kaliyev, Artem Maryanskyy
Multi-agent LLM coordination papers report small benchmark deltas as evidence that one architecture beats another. A prior question: how much paired trial-0 disagreement do two pro…
cs.MA2026
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
Artem Maryanskyy, Dmitry Budnikov, Alibek T. Kaliyev
Multi-agent LLM pipelines produce contradictory evidence on whether team diversity improves output quality: heterogeneous Mixture-of-Agents teams outperform single models, yet homo…