activity
20232026
collaborators
Showing cs.SEShow all

6 papers · 1 filter

cs.SE2026

How Robustly do LLMs Understand Execution Semantics?

Claudio Spiess, Prem Devanbu, Earl T. Barr

LLMs demonstrate remarkable reasoning capabilities, yet whether they utilize internal world models or rely on sophisticated pattern matching remains open. We study LLMs through the…

cs.SE2026

On LLMs' Internal Representation of Code Correctness

Francisco Ribeiro, Claudio Spiess, Prem Devanbu +1

Despite the effectiveness of large language models (LLMs) for code generation, they often output incorrect code. One reason is that model output probabilities are often not well-co…

cs.SE2026

Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs

Ali Al-Kaswan, Claudio Spiess, Prem Devanbu +2

Large language models are increasingly used for code generation and debugging, but their outputs can still contain bugs, that originate from training data. Distinguishing whether a…

cs.SE2025

Does In-IDE Calibration of Large Language Models work at Scale?

Roham Koohestani, Agnia Sergeyuk, David Gros +4

The introduction of large language models into integrated development environments (IDEs) is revolutionizing software engineering, yet it poses challenges to the usefulness and rel…

cs.SE2024

Calibration and Correctness of Language Models for Code

Claudio Spiess, David Gros, Kunal Suresh Pai +6

Machine learning models are widely used, but can also often be wrong. Users would benefit from a reliable indication of whether a given output from a given model should be trusted,…

cs.SE2023

STraceBERT: Source Code Retrieval using Semantic Application Traces

Claudio Spiess

Software reverse engineering is an essential task in software engineering and security, but it can be a challenging process, especially for adversarial artifacts. To address this c…