12 papers
AI Sandbox: Technical Report
Muhammad Waseem, Md Aidul Islam, Md Nasir Uddin Shuvo +8
Collaborative AI experimentation across industry and academia requires platforms that enable rapid prototyping while preserving controlled access, tenant separation, and transparen…
Engineering a Governance-Aware AI Sandbox: Design, Implementation, and Lessons Learned
Muhammad Waseem, Md Aidul Islam, Md Nasir Uddin Shuvo +8
Collaborative AI experimentation in industry-academia requires environments that support rapid trials while maintaining controlled access, organisational isolation, and traceable w…
TDD Governance for Multi-Agent Code Generation via Prompt Engineering
Tarlan Hasanli, Shahbaz Siddeeq, Bishwash Khanal +3
Large language models (LLMs) accelerate software development but often exhibit instability, non-determinism, and weak adherence to development discipline in unconstrained workflows…
Agentic Frameworks for Reasoning Tasks: An Empirical Study
Zeeshan Rasheed, Abdul Malik Sami, Muhammad Waseem +3
Recent advances in agentic frameworks have enabled AI agents to perform complex reasoning and decision-making. However, evidence comparing their reasoning performance, efficiency,…
LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review
Zeeshan Rasheeda, Muhammad Waseema, Kai-Kristian Kemella +2
Large Language Models (LLMs) have enabled multi-agent systems to perform autonomous code generation for complex tasks. Despite the recent growth in research and industrial applicat…
Towards AI Evaluation in Domain-Specific RAG Systems: The AgriHubi Case Study
Md. Toufique Hasan, Ayman Asad Khan, Mika Saari +2
Large language models show promise for knowledge-intensive domains, yet their use in agriculture is constrained by weak grounding, English-centric training data, and limited real-w…