papers

Publications (11)

cs.SE2024

The Counterfeit Conundrum: Can Code Language Models Grasp the Nuances of Their Incorrect Generations?

Alex Gu, Wen-Ding Li, Naman Jain +4

While language models are increasingly more proficient at code generation, they still frequently generate incorrect programs. Many of these programs are obviously wrong, but others…

cs.SE2025

Refactoring Codebases through Library Design

Ziga Kovacic, Justin T. Chiu, Celine Lee +2

Maintainable and general software allows developers to build robust applications efficiently, yet achieving these qualities often requires refactoring specialized solutions into re…

cs.SE2024

Guess & Sketch: Language Model Guided Transpilation

Celine Lee, Abdulrahman Mahmoud, Michal Kurek +5

Maintaining legacy software requires many software and systems engineering hours. Assembly code programs, which demand low-level control over the computer machine state and have no…

cs.AI2025

Critical Thinking: Which Kinds of Complexity Govern Optimal Reasoning Length?

Celine Lee, Alexander M. Rush, Keyon Vafa

Large language models (LLMs) often benefit from verbalized reasoning at inference time, but it remains unclear which aspects of task difficulty these extra reasoning tokens address…

cs.SE2022

MP-CodeCheck: Evolving Logical Expression Code Anomaly Learning with Iterative Self-Supervision

Urs C. Muff, Celine Lee, Paul Gottschlich +1

Machine programming (MP) is concerned with automating software development. According to studies, software engineers spend upwards of 50% of their development time debugging softwa…

cs.SE2021

Toward Code Generation: A Survey and Lessons from Semantic Parsing

Celine Lee, Justin Gottschlich, Dan Roth

With the growth of natural language processing techniques and demand for improved software engineering efficiency, there is an emerging interest in translating intention from human…

cs.SE2024

Commit0: Library Generation from Scratch

Wenting Zhao, Nan Jiang, Celine Lee +4

With the goal of benchmarking generative systems beyond expert software development ability, we introduce Commit0, a benchmark that challenges AI agents to write libraries from scr…

cs.CL2025

Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guarantees

Ahmed Heakl, Sarim Hashmi, Chaimaa Abi +2

The hardware ecosystem is rapidly evolving, with increasing interest in translating low-level programs across different instruction set architectures (ISAs) in a quick, flexible, a…

cs.CL2023

Mixture of Soft Prompts for Controllable Data Generation

Derek Chen, Celine Lee, Yunan Lu +2

Large language models (LLMs) effectively generate fluent text when the target output follows natural language patterns. However, structured prediction tasks confine the output form…

cs.LG2026

The Efficiency Gap in Byte Modeling

Celine Lee, Jing Nathan Yan, Chen Liang +9

Modern language models have historically relied on two dominant design choices: subword tokenization and autoregressive (AR) ordering. These design decisions bake in priors that di…

math.NT2018

On second order linear sequences of composite numbers

Dan Ismailescu, Adrienne Ko, Celine Lee +1

In this paper we present a new proof of the following 2010 result of Dubickas, Novikas, and Siurys: Let and let be the sequence defined by…