Publications (10)
Summarization-Based Document IDs for Generative Retrieval with Language Models
Haoxin Li, Daniel Cheng, Phillip Keung +2
Generative retrieval (Wang et al., 2022; Tay et al., 2022) is a popular approach for end-to-end document retrieval that directly generates document identifiers given an input query…
Systematic Review of Academic Procrastination Interventions in Computing Higher Education
Daniel Cheng, Oscar Heath, Daniyaal Farooqi +3
Academic procrastination is a persistent challenge in computing education, yet evidence on the effectiveness of course-level interventions remains fragmented across diverse designs…
MQXFA Final Design Report
Giorgio Ambrosio, Kathleen Amm, Mike Anerella +29
The MQXFA Quadrupole magnets will be installed in High Luminosity LHC to form the Q1 and Q3 inner triplet optical elements in front of the interaction points 1 (ATLAS) and 5 (CMS).…
Comparison Study: Glacier Calving Front Delineation in Synthetic Aperture Radar Images With Deep Learning
Nora Gourmelon, Konrad Heidler, Erik Loebel +12
Continuous monitoring of glacier calving fronts is essential for sea level rise projections. This study benchmarks Deep Learning systems for front delineation in Synthetic Aperture…
Multi-line AI-assisted Code Authoring
Omer Dunay, Daniel Cheng, Adam Tait +9
CodeCompose is an AI-assisted code authoring tool powered by large language models (LLMs) that provides inline suggestions to 10's of thousands of developers at Meta. In this paper…
AI-assisted Code Authoring at Scale: Fine-tuning, deploying, and mixed methods evaluation
Vijayaraghavan Murali, Chandra Maddila, Imad Ahmad +6
Generative LLMs have been shown to effectively power AI-based code authoring tools that can suggest entire statements or blocks of code during code authoring. In this paper we pres…
Agentic Program Repair from Test Failures at Scale: A Neuro-symbolic approach with static analysis and test execution feedback
Chandra Maddila, Adam Tait, Claire Chang +21
Aim: With the advent of LLMs, sophisticated agentic program repair has become viable at large organizations with large codebases. In this work, we develop an Engineering Agent that…
Expectations Versus Reality: Evaluating Intrusion Detection Systems in Practice
Jake Hesford, Daniel Cheng, Alan Wan +4
Our paper provides empirical comparisons between recent IDSs to provide an objective comparison between them to help users choose the most appropriate solution based on their requi…
Multicolor -Tilings with High Discrepancy
Henry Chan, Daniel Cheng, Lior Gishboliner +1
We study the minimum degree threshold guaranteeing the existence of -tilings of high discrepancy in any -edge-coloring. Balogh, Csaba, Pluhár and Treglown handl…
NarrowBERT: Accelerating Masked Language Model Pretraining and Inference
Haoxin Li, Phillip Keung, Daniel Cheng +2
Large-scale language model pretraining is a very successful form of self-supervised learning in natural language processing, but it is increasingly expensive to perform as the mode…