Publications (249)
Has My System Prompt Been Used? Large Language Model Prompt Membership Inference
Roman Levin, Valeriia Cherepanova, Abhimanyu Hans +2
Prompt engineering has emerged as a powerful technique for optimizing large language models (LLMs) for specific applications, enabling faster prototyping and improved performance,…
Robbing the Fed: Directly Obtaining Private Data in Federated Learning with Modified Models
Liam Fowl, Jonas Geiping, Wojtek Czaja +2
Federated learning has quickly gained popularity with its promises of increased user privacy and efficiency. Previous works have shown that federated gradient updates contain infor…
Improving Generalization of Transfer Learning Across Domains Using Spatio-Temporal Features in Autonomous Driving
Shivam Akhauri, Laura Zheng, Tom Goldstein +1
Practical learning-based autonomous driving models must be capable of generalizing learned behaviors from simulated to real domains, and from training data to unseen domains with u…
Linear Spectral Estimators and an Application to Phase Retrieval
Ramina Ghods, Andrew S. Lan, Tom Goldstein +1
Phase retrieval refers to the problem of recovering real- or complex-valued vectors from magnitude measurements. The best-known algorithms for this problem are iterative in nature…
Nonlinear 1-Bit Precoding for Massive MU-MIMO with Higher-Order Modulation
Sven Jacobsson, Giuseppe Durisi, Mikael Coldrey +2
Massive multi-user (MU) multiple-input multiple- output (MIMO) is widely believed to be a core technology for the upcoming fifth-generation (5G) wireless communication standards. T…
Biconvex Relaxation for Semidefinite Programming in Computer Vision
Sohil Shah, Abhay Kumar, Carlos Castillo +3
Semidefinite programming is an indispensable tool in computer vision, but general-purpose solvers for semidefinite programs are often too slow and memory intensive for large-scale…
The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
Karthik A. Sankararaman, Soham De, Zheng Xu +2
This paper studies how neural network architecture affects the speed of training. We introduce a simple concept called gradient confusion to help formally analyze this. When gradie…
Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks
Micah Goldblum, Hossein Souri, Renkun Ni +10
Neural network based computer vision systems are typically built on a backbone, a pretrained or randomly initialized feature extractor. Several years ago, the default option was an…
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
Siddharth Singh, Prajwal Singhania, Aditya Ranjan +9
Training and fine-tuning large language models (LLMs) with hundreds of billions to trillions of parameters requires tens of thousands of GPUs, and a highly scalable software stack.…
Democratic Representations
Christoph Studer, Tom Goldstein, Wotao Yin +1
Minimization of the (or maximum) norm subject to a constraint that imposes consistency to an underdetermined system of linear equations finds use in a large number…
PHORECAST: Enabling AI Understanding of Public Health Outreach Across Populations
Rifaa Qadri, Anh Nhat Nhu, Swati Ramnath +6
Understanding how diverse individuals and communities respond to persuasive messaging holds significant potential for advancing personalized and socially aware machine learning. Wh…
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
Ruchit Rawal, Reza Shirkavand, Heng Huang +2
Video large language models have not yet been widely deployed, largely due to their tendency to hallucinate. Typical benchmarks for Video-LLMs rely simply on multiple-choice questi…
Active Learning at the ImageNet Scale
Zeyad Ali Sami Emam, Hong-Min Chu, Ping-Yeh Chiang +4
Active learning (AL) algorithms aim to identify an optimal subset of data for annotation, such that deep neural networks (DNN) can achieve better performance when trained on this l…
Exploring Model Robustness with Adaptive Networks and Improved Adversarial Training
Zheng Xu, Ali Shafahi, Tom Goldstein
Adversarial training has proven to be effective in hardening networks against adversarial examples. However, the gained robustness is limited by network capacity and number of trai…
Decentralized Baseband Processing for Massive MU-MIMO Systems
Kaipeng Li, Rishi Sharan, Yujun Chen +3
Achieving high spectral efficiency in realistic massive multi-user (MU) multiple-input multiple-output (MIMO) wireless systems requires computationally-complex algorithms for data…
A New Rank Constraint on Multi-view Fundamental Matrices, and its Application to Camera Location Recovery
Soumyadip Sengupta, Tal Amir, Meirav Galun +4
Accurate estimation of camera matrices is an important step in structure from motion algorithms. In this paper we introduce a novel rank constraint on collections of fundamental ma…
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
Ang Li, Charles Wang, Deqing Fu +9
Humans often use visual aids, for example diagrams or sketches, when solving complex problems. Training multimodal models to do the same, known as Visual Chain of Thought (Visual C…
Dataset Security for Machine Learning: Data Poisoning, Backdoor Attacks, and Defenses
Micah Goldblum, Dimitris Tsipras, Chulin Xie +6
As machine learning systems grow in scale, so do their training data requirements, forcing practitioners to automate and outsource the curation of training data in order to achieve…
Making an Invisibility Cloak: Real World Adversarial Attacks on Object Detectors
Zuxuan Wu, Ser-Nam Lim, Larry Davis +1
We present a systematic study of adversarial attacks on state-of-the-art object detection frameworks. Using standard detection datasets, we train patterns that suppress the objectn…
Certified Data Removal from Machine Learning Models
Chuan Guo, Tom Goldstein, Awni Hannun +1
Good data stewardship requires removal of data at the request of the data's owner. This raises the question if and how a trained machine-learning model, which implicitly stores inf…
Can Neural Nets Learn the Same Model Twice? Investigating Reproducibility and Double Descent from the Decision Boundary Perspective
Gowthami Somepalli, Liam Fowl, Arpit Bansal +5
We discuss methods for visualizing neural network decision boundaries and decision regions. We use these visualizations to investigate issues related to reproducibility and general…
Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models
Gowthami Somepalli, Vasu Singla, Micah Goldblum +2
Cutting-edge diffusion models produce images with high quality and customizability, enabling them to be used for commercial art and graphic design purposes. But do diffusion models…
Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes
Mayuka Jayawardhana, Renbo, Samuel Dooley +6
Large language models (LLMs) perform remarkably well on tabular datasets in zero- and few-shot settings, since they can extract meaning from natural language column headers that de…
MaxVA: Fast Adaptation of Step Sizes by Maximizing Observed Variance of Gradients
Chen Zhu, Yu Cheng, Zhe Gan +3
Adaptive gradient methods such as RMSProp and Adam use exponential moving estimate of the squared gradient to compute adaptive step sizes, achieving better convergence than SGD in…
Certifying Confidence via Randomized Smoothing
Aounon Kumar, Alexander Levine, Soheil Feizi +1
Randomized smoothing has been shown to provide good certified-robustness guarantees for high-dimensional classification problems. It uses the probabilities of predicting the top tw…
DCAN: Dual Channel-wise Alignment Networks for Unsupervised Scene Adaptation
Zuxuan Wu, Xintong Han, Yen-Liang Lin +4
Harvesting dense pixel-level annotations to train deep neural networks for semantic segmentation is extremely expensive and unwieldy at scale. While learning from synthetic data wh…
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
Yuxin Wen, Arman Zharmagambetov, Ivan Evtimov +4
Prompt injection poses a serious threat to the reliability and safety of LLM agents. Recent defenses against prompt injection, such as Instruction Hierarchy and SecAlign, have show…
End-to-End Context Compression at Scale
Ang Li, Sean McLeish, Haozhe Chen +12
Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degra…
A Field Guide to Forward-Backward Splitting with a FASTA Implementation
Tom Goldstein, Christoph Studer, Richard Baraniuk
Non-differentiable and constrained optimization play a key role in machine learning, signal and image processing, communications, and beyond. For high-dimensional minimization prob…
Insta-RS: Instance-wise Randomized Smoothing for Improved Robustness and Accuracy
Chen Chen, Kezhi Kong, Peihong Yu +3
Randomized smoothing (RS) is an effective and scalable technique for constructing neural network classifiers that are certifiably robust to adversarial perturbations. Most RS works…
JPEG Compressed Images Can Bypass Protections Against AI Editing
Pedro Sandoval-Segura, Jonas Geiping, Tom Goldstein
Recently developed text-to-image diffusion models make it easy to edit or create high-quality images. Their ease of use has raised concerns about the potential for malicious editin…
Where do Models go Wrong? Parameter-Space Saliency Maps for Explainability
Roman Levin, Manli Shu, Eitan Borgnia +3
Conventional saliency maps highlight input features to which neural network predictions are highly sensitive. We take a different approach to saliency, in which we identify and ana…
Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs
Ryan Synk, Monte Hoover, John Kirchenbauer +6
There is growing demand for performing inference with hundreds of thousands of input tokens on trained transformer models. Inference at this extreme scale demands significant compu…
A Cookbook of Self-Supervised Learning
Randall Balestriero, Mark Ibrahim, Vlad Sobal +16
Self-supervised learning, dubbed the dark matter of intelligence, is a promising path to advance machine learning. Yet, much like cooking, training SSL methods is a delicate art wi…
What Doesn't Kill You Makes You Robust(er): How to Adversarially Train against Data Poisoning
Jonas Geiping, Liam Fowl, Gowthami Somepalli +3
Data poisoning is a threat model in which a malicious actor tampers with training data to manipulate outcomes at inference time. A variety of defenses against this threat model hav…
WITCHcraft: Efficient PGD attacks with random step size
Ping-Yeh Chiang, Jonas Geiping, Micah Goldblum +4
State-of-the-art adversarial attacks on neural networks use expensive iterative methods and numerous random restarts from different initial points. Iterative FGSM-based methods wit…
Adversarially Robust Distillation
Micah Goldblum, Liam Fowl, Soheil Feizi +1
Knowledge distillation is effective for producing small, high-performance neural networks for classification, but these small networks are vulnerable to adversarial attacks. This p…
Benchmarking Correctness and Security in Multi-Turn Code Generation
Ruchit Rawal, Jeffrey Yang Fan Chiang, Chihao Shen +4
AI coding assistants powered by large language models (LLMs) have transformed software development, significantly boosting productivity. While existing benchmarks evaluate the corr…
Sleeper Agent: Scalable Hidden Trigger Backdoors for Neural Networks Trained from Scratch
Hossein Souri, Liam Fowl, Rama Chellappa +2
As the curation of data for machine learning becomes increasingly automated, dataset tampering is a mounting threat. Backdoor attackers tamper with training data to embed a vulnera…
Certified Defenses for Adversarial Patches
Ping-Yeh Chiang, Renkun Ni, Ahmed Abdelkader +3
Adversarial patch attacks are among one of the most practical threat models against real-world computer vision systems. This paper studies certified and empirical defenses against…
PhaseMax: Convex Phase Retrieval via Basis Pursuit
Tom Goldstein, Christoph Studer
We consider the recovery of a (real- or complex-valued) signal from magnitude-only measurements, known as phase retrieval. We formulate phase retrieval as a convex optimization pro…
Channel Charting: Locating Users within the Radio Environment using Channel State Information
Christoph Studer, Saïd Medjkouh, Emre GönültaŠ+2
We propose channel charting (CC), a novel framework in which a multi-antenna network element learns a chart of the radio geometry in its surrounding area. The channel chart capture…
Technical Challenges for Training Fair Neural Networks
Valeriia Cherepanova, Vedant Nanda, Micah Goldblum +2
As machine learning algorithms have been widely deployed across applications, many concerns have been raised over the fairness of their predictions, especially in high stakes setti…
Adaptive Primal-Dual Hybrid Gradient Methods for Saddle-Point Problems
Tom Goldstein, Min Li, Xiaoming Yuan +2
The Primal-Dual hybrid gradient (PDHG) method is a powerful optimization scheme that breaks complex problems into simple sub-steps. Unfortunately, PDHG methods require the user to…
Datasets for Studying Generalization from Easy to Hard Examples
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta +5
We describe new datasets for studying generalization from easy to hard examples.
Neural Auctions Compromise Bidder Information
Alex Stein, Avi Schwarzschild, Michael Curry +2
Single-shot auctions are commonly used as a means to sell goods, for example when selling ad space or allocating radio frequencies, however devising mechanisms for auctions with mu…
Truth or Backpropaganda? An Empirical Investigation of Deep Learning Theory
Micah Goldblum, Jonas Geiping, Avi Schwarzschild +2
We empirically evaluate common assumptions about neural networks that are widely held by practitioners and theorists alike. In this work, we: (1) prove the widespread existence of…
Instance adaptive adversarial training: Improved accuracy tradeoffs in neural nets
Yogesh Balaji, Tom Goldstein, Judy Hoffman
Adversarial training is by far the most successful strategy for improving robustness of neural networks to adversarial attacks. Despite its success as a defense mechanism, adversar…
On the Reliability of Watermarks for Large Language Models
John Kirchenbauer, Jonas Geiping, Yuxin Wen +7
As LLMs become commonplace, machine-generated text has the potential to flood the internet with spam, social media bots, and valueless content. Watermarking is a simple and effecti…
GenQA: Generating Millions of Instructions from a Handful of Prompts
Jiuhai Chen, Rifaa Qadri, Yuxin Wen +4
Most public instruction finetuning datasets are relatively small compared to the closed source datasets used to train industry models. To study questions about finetuning at scale,…
Identifying and Evaluating Inactive Heads in Pretrained LLMs
Pedro Sandoval-Segura, Xijun Wang, Ashwinee Panda +4
Attention is foundational to large language models (LLMs), enabling different heads to have diverse focus on relevant input tokens. However, learned behaviors like attention sinks,…
PhasePack: A Phase Retrieval Library
Rohan Chandra, Ziyuan Zhong, Justin Hontz +3
Phase retrieval deals with the estimation of complex-valued signals solely from the magnitudes of linear measurements. While there has been a recent explosion in the development of…
Center Smoothing: Certified Robustness for Networks with Structured Outputs
Aounon Kumar, Tom Goldstein
The study of provable adversarial robustness has mostly been limited to classification tasks and models with one-dimensional real-valued outputs. We extend the scope of certifiable…
What Can We Learn from Unlearnable Datasets?
Pedro Sandoval-Segura, Vasu Singla, Jonas Geiping +2
In an era of widespread web scraping, unlearnable dataset methods have the potential to protect data privacy by preventing deep neural networks from generalizing. But in addition t…
Data Augmentation for Meta-Learning
Renkun Ni, Micah Goldblum, Amr Sharaf +2
Conventional image classifiers are trained by randomly sampling mini-batches of images. To achieve state-of-the-art performance, practitioners use sophisticated data augmentation s…
Robustness Disparities in Face Detection
Samuel Dooley, George Z. Wei, Tom Goldstein +1
Facial analysis systems have been deployed by large companies and critiqued by scholars and activists for the past decade. Many existing algorithmic audits examine the performance…
Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery
Yuxin Wen, Neel Jain, John Kirchenbauer +3
The strength of modern generative models lies in their ability to be controlled through text-based prompts. Typical "hard" prompts are made from interpretable words and tokens, and…
Transfer Learning with Deep Tabular Models
Roman Levin, Valeriia Cherepanova, Avi Schwarzschild +5
Recent work on deep learning for tabular data demonstrates the strong performance of deep tabular models, often bridging the gap between gradient boosted decision trees and neural…
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
Juzheng Zhang, Jiacheng You, Ashwinee Panda +1
Low-Rank Adaptation (LoRA) has emerged as a popular parameter-efficient fine-tuning (PEFT) method for Large Language Models (LLMs), yet it still incurs notable overhead and suffers…
A Performance-Driven Benchmark for Feature Selection in Tabular Deep Learning
Valeriia Cherepanova, Roman Levin, Gowthami Somepalli +5
Academic tabular benchmarks often contain small sets of curated features. In contrast, data scientists typically collect as many features as possible into their datasets, and even…
Execute Order 66: Targeted Data Poisoning for Reinforcement Learning
Harrison Foley, Liam Fowl, Tom Goldstein +1
Data poisoning for reinforcement learning has historically focused on general performance degradation, and targeted attacks have been successful via perturbations that involve cont…
A Watermark for Large Language Models
John Kirchenbauer, Jonas Geiping, Yuxin Wen +3
Potential harms of large language models can be mitigated by watermarking model output, i.e., embedding signals into generated text that are invisible to humans but algorithmically…
ODIN: Disentangled Reward Mitigates Hacking in RLHF
Lichang Chen, Chen Zhu, Davit Soselia +6
In this work, we study the issue of reward hacking on the response length, a challenge emerging in Reinforcement Learning from Human Feedback (RLHF) on LLMs. A well-formatted, verb…
VLSI Design of a 3-bit Constant-Modulus Precoder for Massive MU-MIMO
Oscar Castañeda, Sven Jacobsson, Giuseppe Durisi +2
Fifth-generation (5G) cellular systems will build on massive multi-user (MU) multiple-input multiple-output (MIMO) technology to attain high spectral efficiency. However, having hu…
Adversarially robust transfer learning
Ali Shafahi, Parsa Saadatpanah, Chen Zhu +4
Transfer learning, in which a network is trained on one task and re-purposed on another, is often used to produce neural network classifiers when data is scarce or full-scale train…
VLSI Designs for Joint Channel Estimation and Data Detection in Large SIMO Wireless Systems
Oscar Castañeda, Tom Goldstein, Christoph Studer
Channel estimation errors have a critical impact on the reliability of wireless communication systems. While virtually all existing wireless receivers separate channel estimation f…
Comparing Human and Machine Bias in Face Recognition
Samuel Dooley, Ryan Downing, George Wei +10
Much recent research has uncovered and discussed serious concerns of bias in facial analysis technologies, finding performance disparities between groups of people based on perceiv…
Efficient Distributed SGD with Variance Reduction
Soham De, Tom Goldstein
Stochastic Gradient Descent (SGD) has become one of the most popular optimization methods for training machine learning models on massive datasets. However, SGD suffers from two ma…
Curse of Dimensionality on Randomized Smoothing for Certifiable Robustness
Aounon Kumar, Alexander Levine, Tom Goldstein +1
Randomized smoothing, using just a simple isotropic Gaussian distribution, has been shown to produce good robustness guarantees against -norm bounded adversaries. In this w…
Breaking certified defenses: Semantic adversarial examples with spoofed robustness certificates
Amin Ghiasi, Ali Shafahi, Tom Goldstein
To deflect adversarial attacks, a range of "certified" classifiers have been proposed. In addition to labeling an image, certified classifiers produce (when possible) a certificate…
Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Neel Jain, Avi Schwarzschild, Yuxin Wen +7
As Large Language Models quickly become ubiquitous, it becomes critical to understand their security vulnerabilities. Recent work shows that text optimizers can produce jailbreakin…
A Simple and Efficient Baseline for Data Attribution on Images
Vasu Singla, Pedro Sandoval-Segura, Micah Goldblum +2
Data attribution methods play a crucial role in understanding machine learning models, providing insight into which training data points are most responsible for model outputs duri…
Optical image-based thickness characterization of atomically thin nanomaterials using computer vision techniques
Daniel Cui, Tom Goldstein, Jun Yan
The main objective of this study was to develop a novel method of characterizing nanomaterials based on the number of layers without the aid of state-of-the-art electron and force…
Quantized Precoding for Massive MU-MIMO
Sven Jacobsson, Giuseppe Durisi, Mikael Coldrey +2
Massive multiuser (MU) multiple-input multiple-output (MIMO) is foreseen to be one of the key technologies in fifth-generation wireless communication systems. In this paper, we inv…
Image Generation with a Sphere Encoder
Kaiyu Yue, Menglin Jia, Ji Hou +1
We introduce the Sphere Encoder, an efficient generative framework capable of producing images in a single forward pass and competing with many-step diffusion models using fewer th…
DynaGuard: A Dynamic Guardian Model With User-Defined Policies
Monte Hoover, Vatsal Baherwani, Neel Jain +7
Guardian models play a crucial role in ensuring the safety and ethical behavior of user-facing AI applications by enforcing guardrails and detecting harmful content. While standard…
The CLRS-Text Algorithmic Reasoning Language Benchmark
Larisa Markeeva, Sean McLeish, Borja Ibarz +7
Eliciting reasoning capabilities from language models (LMs) is a critical direction on the path towards building intelligent systems. Most recent studies dedicated to reasoning foc…
Transferable Clean-Label Poisoning Attacks on Deep Neural Nets
Chen Zhu, W. Ronny Huang, Ali Shafahi +4
Clean-label poisoning attacks inject innocuous looking (and "correctly" labeled) poison images into training data, causing a model to misclassify a targeted image after being train…
ChannelTok: Efficient Flexible-Length Vision Tokenization
Sukriti Paul, Arpit Bansal, Tom Goldstein
Leading flexible vision tokenizers achieve SOTA quality at an extreme cost, relying on parameter-heavy backbones and slow, multi-step generative decoders. We depart from this compl…
The Intrinsic Dimension of Images and Its Impact on Learning
Phillip Pope, Chen Zhu, Ahmed Abdelkader +2
It is widely believed that natural image data exhibits low-dimensional structure despite the high dimensionality of conventional pixel representations. This idea underlies a common…
Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks
Ali Shafahi, W. Ronny Huang, Mahyar Najibi +4
Data poisoning is an attack on machine learning models wherein the attacker adds examples to the training set to manipulate the behavior of the model at test time. This paper explo…
Son of Zorn's Lemma: Targeted Style Transfer Using Instance-aware Semantic Segmentation
Carlos Castillo, Soham De, Xintong Han +3
Style transfer is an important task in which the style of a source image is mapped onto that of a target image. The method is useful for synthesizing derivative works of a particul…
SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training
Gowthami Somepalli, Micah Goldblum, Avi Schwarzschild +2
Tabular data underpins numerous high-impact applications of machine learning from fraud detection to genomics and healthcare. Classical approaches to solving tabular problems, such…
Finite-Alphabet Wiener Filter Precoding for mmWave Massive MU-MIMO Systems
Oscar Castañeda, Sven Jacobsson, Giuseppe Durisi +2
Power consumption of multi-user (MU) precoding is a major concern in all-digital massive MU multiple-input multiple-output (MIMO) base-stations with hundreds of antenna elements op…
Are adversarial examples inevitable?
Ali Shafahi, W. Ronny Huang, Christoph Studer +2
A wide range of defenses have been proposed to harden neural networks against adversarial attacks. However, a pattern has emerged in which the majority of adversarial defenses are…
EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM
Quang Nguyen, Truong Vu, Trong-Tung Nguyen +6
Image editing technologies are tools used to transform, adjust, remove, or otherwise alter images. Recent research has significantly improved the capabilities of image editing tool…
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
Xiyao Wang, Jiuhai Chen, Zhaoyang Wang +8
Large vision-language models (LVLMs) have achieved impressive results in visual question-answering and reasoning tasks through vision instruction tuning on specific datasets. Howev…
VQ-GNN: A Universal Framework to Scale up Graph Neural Networks using Vector Quantization
Mucong Ding, Kezhi Kong, Jingling Li +4
Most state-of-the-art Graph Neural Networks (GNNs) can be defined as a form of graph convolution which can be realized by message passing between direct neighbors or beyond. To sca…
Preventing Unauthorized Use of Proprietary Data: Poisoning for Secure Dataset Release
Liam Fowl, Ping-yeh Chiang, Micah Goldblum +4
Large organizations such as social media companies continually release data, for example user images. At the same time, these organizations leverage their massive corpora of releas…
Adaptive Relaxed ADMM: Convergence Theory and Practical Implementation
Zheng Xu, Mario A. T. Figueiredo, Xiaoming Yuan +2
Many modern computer vision and machine learning applications rely on solving difficult optimization problems that involve non-differentiable objective functions and constraints. T…
Training Quantized Nets: A Deeper Understanding
Hao Li, Soham De, Zheng Xu +3
Currently, deep neural networks are deployed on low-power portable devices by first training a full-precision model using powerful hardware, and then deriving a corresponding low-p…
ProportionNet: Balancing Fairness and Revenue for Auction Design with Deep Learning
Kevin Kuo, Anthony Ostuni, Elizabeth Horishny +5
The design of revenue-maximizing auctions with strong incentive guarantees is a core concern of economic theory. Computational auctions enable online advertising, sourcing, spectru…
Adaptive ADMM with Spectral Penalty Parameter Selection
Zheng Xu, Mario A. T. Figueiredo, Tom Goldstein
The alternating direction method of multipliers (ADMM) is a versatile tool for solving a wide range of constrained optimization problems, with differentiable or non-differentiable…
Cramming: Training a Language Model on a Single GPU in One Day
Jonas Geiping, Tom Goldstein
Recent trends in language modeling have focused on increasing performance through scaling, and have resulted in an environment where training language models is out of reach for mo…
Network Deconvolution
Chengxi Ye, Matthew Evanusa, Hua He +5
Convolution is a central operation in Convolutional Neural Networks (CNNs), which applies a kernel to overlapping regions shifted across the image. However, because of the strong c…
MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning
Zikui Cai, Andrew Wang, Anirudh Satheesh +10
Despite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions. First, they overwhelmingly rely on static i…
Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise
Arpit Bansal, Eitan Borgnia, Hong-Min Chu +6
Standard diffusion models involve an image transform -- adding Gaussian noise -- and an image restoration operator that inverts this degradation. We observe that the generative beh…
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova +5
Detecting text generated by modern large language models is thought to be hard, as both LLMs and humans can exhibit a wide range of complex behaviors. However, we find that a score…
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models
Ruchit Rawal, Reza Shirkavand, Sayak Paul +5
Inference-time scaling for text-to-image generation has progressed from simple Best-of- (BoN) sampling to guided search methods that verify and steer candidate trajectories at i…
Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions
Keita Teranishi, Harshitha Menon, William F. Godoy +25
We discuss the challenges and propose research directions for using AI to revolutionize the development of high-performance computing (HPC) software. AI technologies, in particular…