collaborators

22 papers

cs.SE2026

Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

Balkrishna Giri, Md Toufique Hasan, Jussi Rasku +2

Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Rel…

cs.SE2026

CodeAssay: A Multi-Metric Benchmark with Audited Ground Truth for LLM Code Generation

Shahbaz Siddeeq, Muhammad Waseem, Umar Subhan Malhi +1

Large Language Models are increasingly evaluated for code generation using test-based benchmarks. The validity of such evaluations depends on the reliability of their references an…

cs.SE2026

AI Sandbox: Technical Report

Muhammad Waseem, Md Aidul Islam, Md Nasir Uddin Shuvo +8

Collaborative AI experimentation across industry and academia requires platforms that enable rapid prototyping while preserving controlled access, tenant separation, and transparen…

cs.SE2026

Vibe Coding in Software Development: A Multivocal Literature Review

Shahbaz Siddeeq, Muhammad Waseem, Kai-Kristian Kemell +3

Vibe coding is a software development practice in which developers state intent in natural language and large language models generate code. It is often framed as one-shot promptin…

cs.SE2026

Identifying and Prioritizing Generative AI Use Cases in an Organization: An Industrial Case Study

Malik Abdul Sami, Zeeshan Rasheed, Meri Olenius +4

Organisations are examining how generative AI can support their operational work and decision-making processes. This study investigates how employees in a energy company understand…

cs.SE2026

Engineering a Governance-Aware AI Sandbox: Design, Implementation, and Lessons Learned

Muhammad Waseem, Md Aidul Islam, Md Nasir Uddin Shuvo +8

Collaborative AI experimentation in industry-academia requires environments that support rapid trials while maintaining controlled access, organisational isolation, and traceable w…