papers

Publications (15)

cs.SE2026

ScarfBench: A Benchmark for Cross-Framework Application Migration in Enterprise Java

Advait Pavuluri, Bridget McGinn, Ashita Saxena +6

Java remains central to enterprise software, and many applications outlive their original architecture. Migrating them across frameworks is a behavior-preserving refactoring spanni…

cs.CV2021

NASTransfer: Analyzing Architecture Transferability in Large Scale Neural Architecture Search

Rameswar Panda, Michele Merler, Mayoore Jaiswal +8

Neural Architecture Search (NAS) is an open and challenging problem in machine learning. While NAS offers great promise, the prohibitive computational demand of most of the existin…

cs.CL2023

CoSiNES: Contrastive Siamese Network for Entity Standardization

Jiaqing Yuan, Michele Merler, Mihir Choudhury +3

Entity standardization maps noisy mentions from free-form text to standard entities in a knowledge base. The unique challenge of this task relative to other entity-related tasks is…

cs.SE2026

Usage, Effects and Requirements for AI Coding Assistants in the Enterprise: An Empirical Study

Maja Vukovic, Rangeet Pan, Tin Kam Ho +3

The rise of large language models (LLMs) has accelerated the development of automated techniques and tools for supporting various software engineering tasks, e.g., program understa…

cs.AI2024

Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Mayank Mishra, Matt Stallone, Gaoyuan Zhang +43

Large Language Models (LLMs) trained on code are revolutionizing the software development process. Increasingly, code LLMs are being integrated into software development environmen…

cs.CL2026

Multi-task Code LLMs: Data Mix or Model Merge?

Mingzhi Zhu, Boris Sobolev, Rahul Krishna +3

Recent research advocates deploying smaller, specialized code LLMs in agentic frameworks alongside frontier models, sparking interest in efficient strategies for multi-task learnin…

eess.IV2020

Covering the News with (AI) Style

Michele Merler, Cicero Nogueira dos Santos, Mauro Martino +2

We introduce a multi-modal discriminative and generative frame-work capable of assisting humans in producing visual content re-lated to a given theme, starting from a collection of…

cs.CV2017

Automatic Curation of Golf Highlights using Multimodal Excitement Features

Michele Merler, Dhiraj Joshi, Quoc-Bao Nguyen +4

The production of sports highlight packages summarizing a game's most exciting moments is an essential task for broadcast media. Yet, it requires labor-intensive video editing. We…

cs.CV2020

Large Scale Neural Architecture Search with Polyharmonic Splines

Ulrich Finkler, Michele Merler, Rameswar Panda +8

Neural Architecture Search (NAS) is a powerful tool to automatically design deep neural networks for many tasks, including image classification. Due to the significant computationa…

cs.SE2026

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing

Mingzhi Zhu, Michele Merler, Raju Pavuluri +1

Code agents must both reason over long-horizon repository state and obey strict tool-use protocols. In paired Instruct/Thinking checkpoints, these capabilities are complementary bu…

cs.AI2025

Toward Continuous Neurocognitive Monitoring: Integrating Speech AI with Relational Graph Transformers for Rare Neurological Diseases

Raquel Norel, Michele Merler, Pavitra Modi

Patients with rare neurological diseases report cognitive symptoms -"brain fog"- invisible to traditional tests. We propose continuous neurocognitive monitoring via smartphone spee…

cs.CL2023

A Comparative Analysis of Task-Agnostic Distillation Methods for Compressing Transformer Language Models

Takuma Udagawa, Aashka Trivedi, Michele Merler +1

Large language models have become a vital component in modern NLP, achieving state of the art performance in a variety of tasks. However, they are often inefficient for real-world…

cs.CL2023

Neural Architecture Search for Effective Teacher-Student Knowledge Transfer in Language Models

Aashka Trivedi, Takuma Udagawa, Michele Merler +3

Large pretrained language models have achieved state-of-the-art results on a variety of downstream tasks. Knowledge Distillation (KD) into a smaller student model addresses their i…

cs.SE2024

Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code

Rangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna +7

Code translation aims to convert source code from one programming language (PL) to another. Given the promising abilities of large language models (LLMs) in code synthesis, researc…

cs.CV2019

Diversity in Faces

Michele Merler, Nalini Ratha, Rogerio S. Feris +1

Face recognition is a long standing challenge in the field of Artificial Intelligence (AI). The goal is to create systems that accurately detect, recognize, verify, and understand…