1 citations · 1 across the 2 of their papers we have counts for
5 papers
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
Ali Al-Kaswan, Maksim Plotnikov, Maxim Hájek +3
Large Language Model (LLM) agents are increasingly proposed for autonomous cybersecurity tasks, but their capabilities in realistic offensive settings remain poorly understood. We…
Learning from Change: Predictive Models for Incident Prevention in a Regulated IT Environment
Eileen Kapel, Jan Lennartz, Luis Cruz +2
Effective IT change management is important for businesses that depend on software and services, particularly in highly regulated sectors such as finance, where operational reliabi…
WaveStitch: Flexible and Fast Conditional Time Series Generation with Diffusion Models
Aditya Shankar, Lydia Y. Chen, Arie van Deursen +1
Generating temporal data under conditions is crucial for forecasting, imputation, and generative tasks. Such data often has metadata and partially observed signals that jointly inf…
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
Jonathan Katzy, Yongcheng Huang, Gopal-Raj Panchu +5
Large Language Models are essential coding assistants, yet their training is predominantly English-centric. In this study, we evaluate the performance of code language models in no…
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
Ali Al-Kaswan, Sebastian Deatc, Begüm Koç +2
Nowadays, developers increasingly rely on solutions powered by Large Language Models (LLM) to assist them with their coding tasks. This makes it crucial to align these tools with h…