Publications (57)
Your Classifier is Secretly an Energy Based Model and You Should Treat it Like One
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen +3
Neural Networks for Modeling Source Code Edits
Rui Zhao, David Bieber, Kevin Swersky +1
Flexibly Fair Representation Learning by Disentanglement
Elliot Creager, David Madras, Jörn-Henrik Jacobsen +4
Pre-trained Gaussian Processes for Bayesian Optimization
Zi Wang, George E. Dahl, Kevin Swersky +5
Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples
Eleni Triantafillou, Tyler Zhu, Vincent Dumoulin +8
Prototypical Networks for Few-shot Learning
Jake Snell, Kevin Swersky, Richard S. Zemel
CUF: Continuous Upsampling Filters
Cristina Vasconcelos, Cengiz Oztireli, Mark Matthews +3
Data-Driven Offline Optimization For Architecting Hardware Accelerators
Aviral Kumar, Amir Yazdanbakhsh, Milad Hashemi +2
Freeze-Thaw Bayesian Optimization
Kevin Swersky, Jasper Snoek, Ryan Prescott Adams
High Mutual Information in Representation Learning with Symmetric Variational Inference
Micha Livne, Kevin Swersky, David J. Fleet
Low-Variance Gradient Estimation in Unrolled Computation Graphs with ES-Single
Paul Vicol, Zico Kolter, Kevin Swersky
The Variational Fair Autoencoder
Christos Louizos, Kevin Swersky, Yujia Li +2
Predicting Deep Zero-Shot Convolutional Neural Networks using Textual Descriptions
Jimmy Ba, Kevin Swersky, Sanja Fidler +1
Fast Exact Inference for Recursive Cardinality Models
Daniel Tarlow, Kevin Swersky, Richard S. Zemel +2
Input Warping for Bayesian Optimization of Non-stationary Functions
Jasper Snoek, Kevin Swersky, Richard S. Zemel +1
Oops I Took A Gradient: Scalable Sampling for Discrete Distributions
Will Grathwohl, Kevin Swersky, Milad Hashemi +2
Learning Execution through Neural Code Fusion
Zhan Shi, Kevin Swersky, Daniel Tarlow +2
Raiders of the Lost Architecture: Kernels for Bayesian Optimization in Conditional Parameter Spaces
Kevin Swersky, David Duvenaud, Jasper Snoek +2
Video models are zero-shot learners and reasoners
Thaddäus Wiedemer, Yuxuan Li, Paul Vicol +6
Scalable Bayesian Optimization Using Deep Neural Networks
Jasper Snoek, Oren Rippel, Kevin Swersky +6
A Comparative Analysis of LLM Adaptation: SFT, LoRA, and ICL in Data-Scarce Scenarios
Bernd Bohnet, Rumen Dangovski, Kevin Swersky +4
Learning to Improve Code Efficiency
Binghong Chen, Daniel Tarlow, Kevin Swersky +5
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
Bernd Bohnet, Kevin Swersky, Rosanne Liu +9
Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Avi Singh, John D. Co-Reyes, Rishabh Agarwal +38
No MCMC for me: Amortized sampling for fast and stable training of energy-based models
Will Grathwohl, Jacob Kelly, Milad Hashemi +3
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
Jeffrey Jian Ma, Milad Hashemi, Amir Yazdanbakhsh +5
Two Sides of the Same Coin: Heterophily and Oversmoothing in Graph Convolutional Neural Networks
Yujun Yan, Milad Hashemi, Kevin Swersky +2
Meta-Learning for Semi-Supervised Few-Shot Classification
Mengye Ren, Eleni Triantafillou, Sachin Ravi +5
Training Language Models on the Knowledge Graph: Insights on Hallucinations and Their Detectability
Jiri Hron, Laura Culp, Gamaleldin Elsayed +28
Frontier Language Models are not Robust to Adversarial Arithmetic, or "What do I need to say so you agree 2+2=5?
C. Daniel Freeman, Laura Culp, Aaron Parisi +27
Big Self-Supervised Models are Strong Semi-Supervised Learners
Ting Chen, Simon Kornblith, Kevin Swersky +2
Graph Normalizing Flows
Jenny Liu, Aviral Kumar, Jimmy Ba +2
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
Towards Better Out-of-Distribution Generalization of Neural Algorithmic Reasoning Tasks
Sadegh Mahdavi, Kevin Swersky, Thomas Kipf +3
Exploring and Benchmarking the Planning Capabilities of Large Language Models
Bernd Bohnet, Azade Nova, Aaron T Parisi +6
MIM: Mutual Information Machine
Micha Livne, Kevin Swersky, David J. Fleet
Analysis of Optimality of Large Language Models on Planning Problems
Bernd Bohnet, Michael C. Mozer, Kevin Swersky +4
Visual prompt engineering for video models
Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7
Generative Moment Matching Networks
Yujia Li, Kevin Swersky, Richard Zemel
Learning Memory Access Patterns
Milad Hashemi, Kevin Swersky, Jamie A. Smith +5
Estimating the Hessian by Back-propagating Curvature
James Martens, Ilya Sutskever, Kevin Swersky
Neural Execution Engines: Learning to Execute Subroutines
Yujun Yan, Kevin Swersky, Danai Koutra +2
Learned Hardware/Software Co-Design of Neural Accelerators
Zhan Shi, Chirag Sakhuja, Milad Hashemi +2
Directly Fine-Tuning Diffusion Models on Differentiable Rewards
Kevin Clark, Paul Vicol, Kevin Swersky +1
Learning unbiased features
Yujia Li, Kevin Swersky, Richard Zemel
Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
Cristina N. Vasconcelos, Abdullah Rashwan, Austin Waters +22
SentenceMIM: A Latent Variable Language Model
Micha Livne, Kevin Swersky, David J. Fleet
Learning Hard Alignments with Variational Inference
Dieterich Lawson, Chung-Cheng Chiu, George Tucker +3
An Imitation Learning Approach for Cache Replacement
Evan Zheran Liu, Milad Hashemi, Kevin Swersky +2
Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
Bernd Bohnet, Pierre-Alexandre Kamienny, Hanie Sedghi +7
Do generative video models understand physical principles?
Saman Motamed, Laura Culp, Kevin Swersky +2
Pre-training helps Bayesian optimization too
Zi Wang, George E. Dahl, Kevin Swersky +6
Apollo: Transferable Architecture Exploration
Amir Yazdanbakhsh, Christof Angermueller, Berkin Akin +7
An online sequence-to-sequence model for noisy speech recognition
Chung-Cheng Chiu, Dieterich Lawson, Yuping Luo +4
Optimizing Long-term Social Welfare in Recommender Systems: A Constrained Matching Approach
Martin Mladenov, Elliot Creager, Omer Ben-Porat +3
Learning Sparse Networks Using Targeted Dropout
Aidan N. Gomez, Ivan Zhang, Siddhartha Rao Kamalakara +4
Human 3D keypoints via spatial uncertainty modeling
Francis Williams, Or Litany, Avneesh Sud +2