papers

Publications (19)

cs.LG2022

AI Gone Astray: Technical Supplement

Janice Yang, Ludvig Karstens, Casey Ross +1

This study is a technical supplement to "AI gone astray: How subtle shifts in patient data send popular algorithms reeling, undermining patient safety." from STAT News, which inves…

cs.CV2026

Open-Ended CT Volume Segmentation with Weak Supervision from Language

Sanjay Subramanian, Junwei Yu, Zirui Wang +5

We introduce a method for training a text-conditioned segmentation model for CT scans, which combines voxel-level supervision with coarse but scalable slice-level supervision from…

cs.LG2022

Syfer: Neural Obfuscation for Private Data Release

Adam Yala, Victor Quach, Homa Esfahanizadeh +5

Balancing privacy and predictive utility remains a central challenge for machine learning in healthcare. In this paper, we develop Syfer, a neural obfuscation method to protect aga…

cs.CV2025

Rethinking Patch Dependence for Masked Autoencoders

Letian Fu, Long Lian, Renhao Wang +6

In this work, we examine the impact of inter-patch dependencies in the decoder of masked autoencoders (MAE) on representation learning. We decompose the decoding mechanism for mask…

cs.CV2026

Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset

Louis Blankemeier, Ashwin Kumar, Joseph Paul Cohen +37

The large volume of abdominal computed tomography (CT) scans coupled with the shortage of radiologists have intensified the need for automated medical image analysis tools. Previou…

cs.CR2021

NeuraCrypt: Hiding Private Health Data via Random Neural Networks for Public Training

Adam Yala, Homa Esfahanizadeh, Rafael G. L. D' Oliveira +6

Balancing the needs of data privacy and predictive utility is a central challenge for machine learning in healthcare. In particular, privacy concerns have led to a dearth of public…

cs.CV2024

LLM-grounded Video Diffusion Models

Long Lian, Baifeng Shi, Adam Yala +2

Text-conditioned diffusion models have emerged as a promising tool for neural video generation. However, current models still struggle with intricate spatiotemporal prompts and oft…

cs.CV2025

Atlas: Multi-Scale Attention Improves Long Context Image Modeling

Kumar Krishna Agrawal, Long Lian, Longchao Liu +6

Efficiently modeling massive images is a long-standing challenge in machine learning. To this end, we introduce Multi-Scale Attention (MSA). MSA relies on two key ideas, (i) multi-…

cs.AI2025

Learning Adaptive Parallel Reasoning with Language Models

Jiayi Pan, Xiuyu Li, Long Lian +6

Scaling inference-time computation has substantially improved the reasoning capabilities of language models. However, existing methods have significant limitations: serialized chai…

cs.CV2026

Stateful Visual Encoders for Vision-Language Models

Zirui Wang, Junwei Yu, Adam Yala +3

Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes. However, in existing open-weight VLMs, vis…

cs.LG2025

Data reuse enables cost-efficient randomized trials of medical AI models

Michael Nercessian, Wenxin Zhang, Alexander Schubert +4

Randomized controlled trials (RCTs) are indispensable for establishing the clinical value of medical artificial-intelligence (AI) tools, yet their high cost and long timelines hind…

cs.LG2026

ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models

Long Lian, Sida Wang, Felix Juefei-Xu +7

Scaling inference-time computation has enabled Large Language Models (LLMs) to achieve strong reasoning performance, but their inherently sequential decoding incurs substantial lat…

cs.LG2023

PEOPL: Characterizing Privately Encoded Open Datasets with Public Labels

Homa Esfahanizadeh, Adam Yala, Rafael G. L. D'Oliveira +8

Allowing organizations to share their data for training of machine learning (ML) models without unintended information leakage is an open problem in practice. A promising technique…

cs.CV2024

LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Long Lian, Boyi Li, Adam Yala +1

Recent advancements in text-to-image diffusion models have yielded impressive results in generating realistic and diverse images. However, these models still struggle with complex…

cs.CV2025

Pillar-0: A New Frontier for Radiology Foundation Models

Kumar Krishna Agrawal, Longchao Liu, Long Lian +11

Radiology plays an integral role in modern medicine, yet rising imaging volumes have far outpaced workforce growth. Foundation models offer a path toward assisting with the full sp…

cs.CV2025

TULIP: Towards Unified Language-Image Pretraining

Zineng Tang, Long Lian, Seun Eisape +6

Despite the recent success of image-text contrastive models like CLIP and SigLIP, these models often struggle with vision-centric tasks that demand high-fidelity image understandin…

cs.CL2016

Improving Information Extraction by Acquiring External Evidence with Reinforcement Learning

Karthik Narasimhan, Adam Yala, Regina Barzilay

Most successful information extraction systems operate with access to a large collection of documents. In this work, we explore the task of acquiring and incorporating external evi…

cs.CV2025

Describe Anything: Detailed Localized Image and Video Captioning

Long Lian, Yifan Ding, Yunhao Ge +8

Generating detailed and accurate descriptions for specific regions in images and videos remains a fundamental challenge for vision-language models. We introduce the Describe Anythi…

cs.CL2024

Conformal Language Modeling

Victor Quach, Adam Fisch, Tal Schuster +4

We propose a novel approach to conformal prediction for generative language models (LMs). Standard conformal prediction produces prediction sets -- in place of single predictions -…