22 papers
Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems
Balkrishna Giri, Md Toufique Hasan, Jussi Rasku +2
Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Rel…
CodeAssay: A Multi-Metric Benchmark with Audited Ground Truth for LLM Code Generation
Shahbaz Siddeeq, Muhammad Waseem, Umar Subhan Malhi +1
Large Language Models are increasingly evaluated for code generation using test-based benchmarks. The validity of such evaluations depends on the reliability of their references an…
AI Sandbox: Technical Report
Muhammad Waseem, Md Aidul Islam, Md Nasir Uddin Shuvo +8
Collaborative AI experimentation across industry and academia requires platforms that enable rapid prototyping while preserving controlled access, tenant separation, and transparen…
Vibe Coding in Software Development: A Multivocal Literature Review
Shahbaz Siddeeq, Muhammad Waseem, Kai-Kristian Kemell +3
Vibe coding is a software development practice in which developers state intent in natural language and large language models generate code. It is often framed as one-shot promptin…
Identifying and Prioritizing Generative AI Use Cases in an Organization: An Industrial Case Study
Malik Abdul Sami, Zeeshan Rasheed, Meri Olenius +4
Organisations are examining how generative AI can support their operational work and decision-making processes. This study investigates how employees in a energy company understand…
Engineering a Governance-Aware AI Sandbox: Design, Implementation, and Lessons Learned
Muhammad Waseem, Md Aidul Islam, Md Nasir Uddin Shuvo +8
Collaborative AI experimentation in industry-academia requires environments that support rapid trials while maintaining controlled access, organisational isolation, and traceable w…