20 papers
CodeAssay: A Multi-Metric Benchmark with Audited Ground Truth for LLM Code Generation
Shahbaz Siddeeq, Muhammad Waseem, Umar Subhan Malhi +1
Large Language Models are increasingly evaluated for code generation using test-based benchmarks. The validity of such evaluations depends on the reliability of their references an…
Vibe Coding in Software Development: A Multivocal Literature Review
Shahbaz Siddeeq, Muhammad Waseem, Kai-Kristian Kemell +3
Vibe coding is a software development practice in which developers state intent in natural language and large language models generate code. It is often framed as one-shot promptin…
Zero-Shot Heart Rate Variability Forecasting from Consumer Wearables Using Time Series Foundation Models
Luukas Peräkylä, Fahad Sohrab, Ville Hautamäki +3
Short-term Heart Rate Variability (HRV) forecasting could provide clinicians with actionable lead time for detecting autonomic dysfunction and adverse cardiac events. Consumer wear…
Identifying and Prioritizing Generative AI Use Cases in an Organization: An Industrial Case Study
Malik Abdul Sami, Zeeshan Rasheed, Meri Olenius +4
Organisations are examining how generative AI can support their operational work and decision-making processes. This study investigates how employees in a energy company understand…
Epic-Organized vs. Requirement-Aligned Gherkin: An Empirical Evaluation of LLM-Based Acceptance Criteria Generation
Shahbaz Siddeeq, Mateen Abbasi, Jussi Rasku +4
Automated authoring of Gherkin Behavior-Driven Development (BDD) acceptance criteria remains a manual bottleneck in requirements engineering. This study investigates whether epic-o…
Anomaly Detection in Smart Power Grids with Graph-Regularized MS-SVDD: a Multimodal Subspace Learning Approach
Thomas Debelle, Fahad Sohrab, Pekka Abrahamsson +1
Anomaly detection in smart power grids is a critical challenge due to the complexity, heterogeneity, and dynamic nature of sensor data streams. Existing one-class classification me…