3 papers
cs.SE2026
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
Syed Yusuf Ahmed, Shiwei Feng, Chanwoo Bae +1
Autonomous AI agents powered by large language models (LLMs) are increasingly deployed in real-world applications, where reliable and robust behavior is critical. However, existing…
cs.SE2025
State-of-the-art Small Language Coder Model: Mify-Coder
Abhinav Parmar, Abhisek Panigrahi, Abhishek Kumar Dwivedi +93
We present Mify-Coder, a 2.5B-parameter code model trained on 4.2T tokens using a compute-optimal strategy built on the Mify-2.5B foundation model. Mify-Coder achieves comparable a…
cs.SE2025
TAI3: Testing Agent Integrity in Interpreting User Intent
Shiwei Feng, Xiangzhe Xu, Xuan Chen +5
LLM agents are increasingly deployed to automate real-world tasks by invoking APIs through natural language instructions. While powerful, they often suffer from misinterpretation o…