papers

Publications (109)

cs.LG2024

Data-Driven Simulator for Mechanical Circulatory Support with Domain Adversarial Neural Process

Sophia Sun, Wenyuan Chen, Zihao Zhou +3

Mechanical Circulatory Support (MCS) devices, implemented as a probabilistic deep sequence model. Existing mechanical simulators for MCS rely on oversimplifying assumptions and are…

cs.LG2026

Manifold-Guided Attention Steering

Ian Li, Kapilesh Guruprasad, Raunak Sengupta +3

Large language models frequently produce errors in reasoning tasks despite possessing the underlying knowledge required for correct reasoning. One possible approach to improve reas…

cs.LG2024

Neural Point Process for Learning Spatiotemporal Event Dynamics

Zihao Zhou, Xingyi Yang, Ryan Rossi +2

Learning the dynamics of spatiotemporal events is a fundamental problem. Neural point processes enhance the expressivity of point process models with deep neural networks. However,…

cs.LG2024

Multi-Fidelity Residual Neural Processes for Scalable Surrogate Modeling

Ruijia Niu, Dongxia Wu, Kai Kim +3

Multi-fidelity surrogate modeling aims to learn an accurate surrogate at the highest fidelity level by combining data from multiple sources. Traditional methods relying on Gaussian…

physics.comp-ph2026

SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs

Yadi Cao, Sicheng Lai, Jiahe Huang +12

Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@…

cs.LG2024

Technical report: Improving the properties of molecules generated by LIMO

Vineet Thumuluri, Peter Eckmann, Michael K. Gilson +1

This technical report investigates variants of the Latent Inceptionism on Molecules (LIMO) framework to improve the properties of generated molecules. We conduct ablative studies o…

cs.LG2026

VICON: Vision In-Context Operator Networks for Multi-Physics Fluid Dynamics Prediction

Yadi Cao, Yuxuan Liu, Liu Yang +3

In-Context Operator Networks (ICONs) have demonstrated the ability to learn operators across diverse partial differential equations using few-shot, in-context learning. However, ex…

cs.CL2024

MORL-Prompt: An Empirical Analysis of Multi-Objective Reinforcement Learning for Discrete Prompt Optimization

Yasaman Jafari, Dheeraj Mekala, Rose Yu +1

RL-based techniques can be employed to search for prompts that, when fed into a target language model, maximize a set of user-specified reward functions. However, in many target ap…

cs.LG2021

Trajectory Prediction using Equivariant Continuous Convolution

Robin Walters, Jinxi Li, Rose Yu

Trajectory prediction is a critical part of many AI applications, for example, the safe operation of autonomous vehicles. However, current methods are prone to making inconsistent…

cs.LG2024

Latent Space Symmetry Discovery

Jianke Yang, Nima Dehmamy, Robin Walters +1

Equivariant neural networks require explicit knowledge of the symmetry group. Automatic symmetry discovery methods aim to relax this constraint and learn invariance and equivarianc…

cs.LG2020

Aortic Pressure Forecasting with Deep Sequence Learning

Eliza Huang, Rui Wang, Uma Chandrasekaran +1

Mean aortic pressure (MAP) is a major determinant of perfusion in all organs systems. The ability to forecast MAP would enhance the ability of physicians to estimate prognosis of t…

cs.AI2026

Think like a Scientist: Physics-guided LLM Agent for Equation Discovery

Jianke Yang, Ohm Venkatachalam, Mohammad Kianezhad +2

Explaining observed phenomena through symbolic, interpretable formulas is a fundamental goal of science. Recently, large language models (LLMs) have emerged as promising tools for…

cs.LG2025

ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models

Veeramakali Vignesh Manivannan, Yasaman Jafari, Srikar Eranky +6

The use of Large Language Models (LLMs) in climate science has recently gained significant attention. However, a critical issue remains: the lack of a comprehensive evaluation fram…

cs.LG2020

Efficient Tensor Decomposition with Boolean Factors

Sung-En Chang, Xun Zheng, Ian E. H. Yen +2

Tensor decomposition has been extensively used as a tool for exploratory analysis. Motivated by neuroscience applications, we study tensor decomposition with Boolean factors. The r…

cs.LG2020

Dynamic Relational Inference in Multi-Agent Trajectories

Ruichao Xiao, Manish Kumar Singh, Rose Yu

Inferring interactions from multi-agent trajectories has broad applications in physics, vision and robotics. Neural relational inference (NRI) is a deep generative model that can r…

cs.AI2022

Predicting the Future of AI with AI: High-quality link prediction in an exponentially growing knowledge network

Mario Krenn, Lorenzo Buffoni, Bruno Coutinho +13

A tool that could suggest new personalized research directions and ideas by taking insights from the scientific literature could significantly accelerate the progress of science. A…

cs.LG2017

Socratic Learning: Augmenting Generative Models to Incorporate Latent Subsets in Training Data

Paroma Varma, Bryan He, Dan Iter +4

A challenge in training discriminative models like neural networks is obtaining enough labeled training data. Recent approaches use generative models to combine weak supervision so…

cs.LG2026

Breaking the Factorization Barrier in Diffusion Language Models

Ian Li, Zilei Shao, Benjie Wang +3

Diffusion language models theoretically allow for efficient parallel generation but are practically hindered by the ``factorization barrier'': the assumption that simultaneously pr…

cs.CL2025

SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA

Haozhou Xu, Dongxia Wu, Matteo Chinazzi +3

Large Language Models (LLMs) show promise in generating long-form scientific explanations that synthesize evidence and connect multiple factors. However, in long-form scientific qu…

cs.LG2025

AtlasD: Automatic Local Symmetry Discovery

Manu Bhat, Jonghyun Park, Jianke Yang +3

Existing symmetry discovery methods predominantly focus on global transformations across the entire system or space, but they fail to consider the symmetries in local neighborhoods…

cs.LG2023

Deep Bayesian Active Learning for Accelerating Stochastic Simulation

Dongxia Wu, Ruijia Niu, Matteo Chinazzi +3

Stochastic simulations such as large-scale, spatiotemporal, age-structured epidemic models are computationally expensive at fine-grained resolution. While deep surrogate models can…

cs.LG2023

Symmetries, flat minima, and the conserved quantities of gradient flow

Bo Zhao, Iordan Ganev, Robin Walters +2

Empirical studies of the loss landscape of deep networks have revealed that many local minima are connected through low-loss valleys. Yet, little is known about the theoretical ori…

cs.LG2026

ToolMol: Evolutionary Agentic Framework for Multi-objective Drug Discovery

Andrew Y. Zhou, Sharvaree Vadgama, Sumanth Varambally +3

Advances in large language models (LLMs) have recently opened new and promising avenues for small-molecule drug discovery. Yet existing LLM-based approaches for molecular generatio…

cs.LG2023

Disentangled Multi-Fidelity Deep Bayesian Active Learning

Dongxia Wu, Ruijia Niu, Matteo Chinazzi +2

To balance quality and cost, various domain areas of science and engineering run simulations at multiple levels of sophistication. Multi-fidelity active learning aims to learn a di…

cs.LG2025

Understanding Mode Connectivity via Parameter Space Symmetry

Bo Zhao, Nima Dehmamy, Robin Walters +1

Neural network minima are often connected by curves along which train and test loss remain nearly constant, a phenomenon known as mode connectivity. While this property has enabled…

cs.LG2025

MF-LAL: Drug Compound Generation Using Multi-Fidelity Latent Space Active Learning

Peter Eckmann, Dongxia Wu, Germano Heinzelmann +2

Current generative models for drug discovery primarily use molecular docking as an oracle to guide the generation of active compounds. However, such models are often not useful in…

cs.LG2025

Symmetry in Neural Network Parameter Spaces

Bo Zhao, Robin Walters, Rose Yu

Modern deep learning models are highly overparameterized, resulting in large sets of parameter configurations that yield the same outputs. A significant portion of this redundancy…

cs.LG2025

Discovering Latent Causal Graphs from Spatiotemporal Data

Kun Wang, Sumanth Varambally, Duncan Watson-Parris +2

Many important phenomena in scientific fields like climate, neuroscience, and epidemiology are naturally represented as spatiotemporal gridded data with complex interactions. Infer…

cs.LG2023

Koopman Neural Forecaster for Time Series with Temporal Distribution Shifts

Rui Wang, Yihe Dong, Sercan Ö. Arik +1

Temporal distributional shifts, with underlying dynamics changing over time, frequently occur in real-world time series and pose a fundamental challenge for deep neural networks (D…

cs.LG2017

Tensor Regression Meets Gaussian Processes

Rose Yu, Guangyu Li, Yan Liu

Low-rank tensor regression, a new model class that learns high-order correlation from data, has recently received considerable attention. At the same time, Gaussian processes (GP)…

cs.LG2021

Traffic Forecasting using Vehicle-to-Vehicle Communication

Steven Wong, Lejun Jiang, Robin Walters +3

We take the first step in using vehicle-to-vehicle (V2V) communication to provide real-time on-board traffic predictions. In order to best utilize real-world V2V communication data…

cs.LG2021

DeepGLEAM: A hybrid mechanistic and deep learning model for COVID-19 forecasting

Dongxia Wu, Liyao Gao, Xinyue Xiong +4

We introduce DeepGLEAM, a hybrid model for COVID-19 forecasting. DeepGLEAM combines a mechanistic stochastic simulation model GLEAM with deep learning. It uses deep learning to lea…

cs.CL2026

Calibrating LLMs with Semantic-level Reward

Fengfei Yu, Ruijia Niu, Dongxia Wu +2

As large language models (LLMs) are deployed in consequential settings such as medical question answering and legal reasoning, the ability to estimate when their outputs are likely…

cs.LG2026

Divide and Learn: Multi-Objective Combinatorial Optimization at Scale

Esha Singh, Dongxia Wu, Chien-Yi Yang +3

Multi-objective combinatorial optimization seeks Pareto-optimal solutions over exponentially large discrete spaces, yet existing methods sacrifice generality, scalability, or theor…

cs.LG2025

Artificial Intelligence for Science in Quantum, Atomistic, and Continuum Systems

Xuan Zhang, Limei Wang, Jacob Helwig +60

Advances in artificial intelligence (AI) are fueling a new paradigm of discoveries in natural sciences. Today, AI has started to advance natural sciences by improving, accelerating…

physics.plasm-ph2026

TGLF-WINN: Data-Efficient Deep Learning Surrogate for Turbulent Transport Modeling in Fusion

Yadi Cao, Futian Zhang, Wesley Liu +7

The Trapped Gyro-Landau Fluid (TGLF) model provides fast, accurate predictions of turbulent transport in tokamaks, but whole device simulations requiring thousands of evaluations r…

stat.ML2024

Long-term Forecasting with TiDE: Time-series Dense Encoder

Abhimanyu Das, Weihao Kong, Andrew Leach +3

Recent work has shown that simple linear models can outperform several Transformer based approaches in long term time-series forecasting. Motivated by this, we propose a Multi-laye…

cs.LG2021

Incorporating Symmetry into Deep Dynamics Models for Improved Generalization

Rui Wang, Robin Walters, Rose Yu

Recent work has shown deep learning can accelerate the prediction of physical dynamics relative to numerical solvers. However, limited physical accuracy and an inability to general…

cs.LG2023

Probabilistic Symmetry for Multi-Agent Dynamics

Sophia Sun, Robin Walters, Jinxi Li +1

Learning multi-agent dynamics is a core AI problem with broad applications in robotics and autonomous driving. While most existing works focus on deterministic prediction, producin…

cs.AI2026

Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation Spaces

Pratham Yashwante, Rose Yu

The Platonic Representation Hypothesis posits that learned representations from models trained on different modalities converge to a shared latent structure of the world. However,…

cs.LG2024

Probabilistic Emulation of a Global Climate Model with Spherical DYffusion

Salva Rühling Cachay, Brian Henn, Oliver Watt-Meyer +2

Data-driven deep learning models are transforming global weather forecasting. It is an open question if this success can extend to climate modeling, where the complexity of the dat…

cs.LG2024

Learning Granger Causality from Instance-wise Self-attentive Hawkes Processes

Dongxia Wu, Tsuyoshi Idé, Aurélie Lozano +5

We address the problem of learning Granger causality from asynchronous, interdependent, multi-type event sequences. In particular, we are interested in discovering instance-level c…

cs.LG2023

Target-Free Compound Activity Prediction via Few-Shot Learning

Peter Eckmann, Jake Anderson, Michael K. Gilson +1

Predicting the activities of compounds against protein-based or phenotypic assays using only a few known compounds and their activities is a common task in target-free drug discove…

cs.AI2026

Hilbert: Recursively Building Formal Proofs with Informal Reasoning

Sumanth Varambally, Thomas Voice, Yanchao Sun +3

Large Language Models (LLMs) demonstrate impressive mathematical reasoning abilities, but their solutions frequently contain errors that cannot be automatically checked. Formal the…

cs.LG2024

Back to Bayesics: Uncovering Human Mobility Distributions and Anomalies with an Integrated Statistical and Neural Framework

Minxuan Duan, Yinlong Qian, Lingyi Zhao +4

Existing methods for anomaly detection often fall short due to their inability to handle the complexity, heterogeneity, and high dimensionality inherent in real-world mobility data…

cs.LG2026

Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs

Ruijia Niu, Dongxia Wu, Rose Yu +1

Accurate uncertainty quantification in large language models (LLMs) is essential for reliable confidence estimation, yet fine-tuned LLMs often become overconfident under limited ad…

q-bio.QM2022

AI-Bind: Improving Binding Predictions for Novel Protein Targets and Ligands

Ayan Chatterjee, Robin Walters, Zohair Shafi +7

Identifying novel drug-target interactions (DTI) is a critical and rate limiting step in drug discovery. While deep learning models have been proposed to accelerate the identificat…

cs.LG2021

Bridging Physics-based and Data-driven modeling for Learning Dynamical Systems

Rui Wang, Danielle Maddix, Christos Faloutsos +2

How can we learn a dynamical system to make forecasts, when some variables are unobserved? For instance, in COVID-19, we want to forecast the number of infected and death cases but…

cs.LG2026

Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution

Ning Liu, Kalle Kujanpää, Zhaoxuan Zhu +11

Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constr…

cs.LG2022

Faster Optimization on Sparse Graphs via Neural Reparametrization

Nima Dehmamy, Csaba Both, Jianzhi Long +1

In mathematical optimization, second-order Newton's methods generally converge faster than first-order methods, but they require the inverse of the Hessian, hence are computational…

cs.LG2020

Learning Disentangled Representations of Video with Missing Data

Armand Comas-Massagué, Chi Zhang, Zlatan Feric +2

Missing data poses significant challenges while learning representations of video sequences. We present Disentangled Imputed Video autoEncoder (DIVE), a deep generative model that…

cs.SI2020

Finding Patient Zero: Learning Contagion Source with Graph Neural Networks

Chintan Shah, Nima Dehmamy, Nicola Perra +4

Locating the source of an epidemic, or patient zero (P0), can provide critical insights into the infection's transmission course and allow efficient resource allocation. Existing m…

q-bio.BM2024

MFBind: a Multi-Fidelity Approach for Evaluating Drug Compounds in Practical Generative Modeling

Peter Eckmann, Dongxia Wu, Germano Heinzelmann +2

Current generative models for drug discovery primarily use molecular docking to evaluate the quality of generated compounds. However, such models are often not useful in practice b…

cs.LG2016

Learning from Multiway Data: Simple and Efficient Tensor Regression

Rose Yu, Yan Liu

Tensor regression has shown to be advantageous in learning tasks with multi-directional relatedness. Given massive multiway data, traditional methods are often too slow to operate…

physics.chem-ph2022

SELFIES and the future of molecular string representations

Mario Krenn, Qianxiang Ai, Senja Barthel +28

Artificial intelligence (AI) and machine learning (ML) are expanding in popularity for broad applications to challenging tasks in chemistry and materials science. Examples include…

cs.CV2021

Generator Surgery for Compressed Sensing

Niklas Smedemark-Margulies, Jung Yeon Park, Max Daniels +3

Image recovery from compressive measurements requires a signal prior for the images being reconstructed. Recent work has explored the use of deep generative models with low latent…

cs.RO2020

Deep Imitation Learning for Bimanual Robotic Manipulation

Fan Xie, Alexander Chowdhury, M. Clara De Paolis Kaluza +3

We present a deep imitation learning framework for robotic bimanual manipulation in a continuous state-action space. A core challenge is to generalize the manipulation skills to ob…

cs.LG2024

Understanding the Difficulty of Solving Cauchy Problems with PINNs

Tao Wang, Bo Zhao, Sicun Gao +1

Physics-Informed Neural Networks (PINNs) have gained popularity in scientific computing in recent years. However, they often fail to achieve the same level of accuracy as classical…

cs.LG2026

Assessing Low Back Movement with Motion Tape Sensor Data Through Deep Learning

Jared Levy, Aarti Lalwani, Elijah Wyckoff +4

Back pain is a pervasive issue affecting a significant portion of the population, often worsened by certain movements of the lower back. Assessing these movements is important for…

cs.LG2026

CaTS-Bench: Can Language Models Describe Time Series?

Luca Zhou, Pratham Yashwante, Marshall Fisher +4

Time series captioning, the task of describing time series in natural language, requires numeric and temporal reasoning, trend interpretation, and contextual understanding. Existin…

cs.AI2021

Quantifying Uncertainty in Deep Spatiotemporal Forecasting

Dongxia Wu, Liyao Gao, Xinyue Xiong +4

Deep learning is gaining increasing popularity for spatiotemporal forecasting. However, prior works have mostly focused on point estimates without quantifying the uncertainty of th…

physics.comp-ph2020

Towards Physics-informed Deep Learning for Turbulent Flow Prediction

Rui Wang, Karthik Kashinath, Mustafa Mustafa +2

While deep learning has shown tremendous success in a wide range of domains, it remains a grand challenge to incorporate physical principles in a systematic manner to the design, t…

cs.AI2025

Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025

Nitya Thakkar, Mert Yuksekgonul, Jake Silberg +6

Peer review at AI conferences is stressed by rapidly rising submission volumes, leading to deteriorating review quality and increased author dissatisfaction. To address these issue…

cs.LG2022

Multi-fidelity Hierarchical Neural Processes

Dongxia Wu, Matteo Chinazzi, Alessandro Vespignani +2

Science and engineering fields use computer simulation extensively. These simulations are often run at multiple levels of sophistication to balance accuracy and efficiency. Multi-f…

cs.LG2025

Conformal Prediction for Time-series Forecasting with Change Points

Sophia Sun, Rose Yu

Conformal prediction has been explored as a general and efficient way to provide uncertainty quantification for time series. However, current methods struggle to handle time series…

cs.LG2022

Taming the Long Tail of Deep Probabilistic Forecasting

Jedrzej Kozerawski, Mayank Sharan, Rose Yu

Deep probabilistic forecasting is gaining attention in numerous applications ranging from weather prognosis, through electricity consumption estimation, to autonomous vehicle traje…

cs.LG2023

Understanding why shooters shoot -- An AI-powered engine for basketball performance profiling

Alejandro Rodriguez Pascual, Ishan Mehta, Muhammad Khan +2

Understanding player shooting profiles is an essential part of basketball analysis: knowing where certain opposing players like to shoot from can help coaches neutralize offensive…

cs.LG2019

NAOMI: Non-Autoregressive Multiresolution Sequence Imputation

Yukai Liu, Rose Yu, Stephan Zheng +2

Missing value imputation is a fundamental problem in spatiotemporal modeling, from motion tracking to the dynamics of physical systems. Deep autoregressive models suffer from error…

cs.LG2026

A Survey of Weight Space Learning: Understanding, Representation, and Generation

Xiaolong Han, Zehong Wang, Bo Zhao +8

Neural network weights are typically viewed as the end product of training, while most deep learning research focuses on data, features, and architectures. However, recent advances…

cs.LG2024

Symmetry-Informed Governing Equation Discovery

Jianke Yang, Wang Rao, Nima Dehmamy +2

Despite the advancements in learning governing differential equations from observations of dynamical systems, data-driven methods are often unaware of fundamental physical laws, su…

cs.LG2024

Copula Conformal Prediction for Multi-step Time Series Forecasting

Sophia Sun, Rose Yu

Accurate uncertainty measurement is a key step to building robust and reliable machine learning systems. Conformal prediction is a distribution-free uncertainty quantification algo…

cs.LG2025

Guardian-regularized Safe Offline Reinforcement Learning for Smart Weaning of Mechanical Circulatory Devices

Aysin Tumay, Sophia Sun, Sonia Fereidooni +3

We study the sequential decision-making problem for automated weaning of mechanical circulatory support (MCS) devices in cardiogenic shock patients. MCS devices are percutaneous mi…

cs.RO2019

Neural Lander: Stable Drone Landing Control using Learned Dynamics

Guanya Shi, Xichen Shi, Michael O'Connell +5

Precise near-ground trajectory control is difficult for multi-rotor drones, due to the complex aerodynamic effects caused by interactions between multi-rotor airflow and the enviro…

cs.LG2026

Generative OOD-regularized Model-based Policy Optimization

Aysin Tumay, Jiahe Huang, Elise Jortberg +1

We study sequential decision-making with offline reinforcement learning (RL). Traditional offline RL policies may result in out-of-distribution (OOD) actions when training relies o…

cs.LG2025

Elucidated Rolling Diffusion Models for Probabilistic Forecasting of Complex Dynamics

Salva Rühling Cachay, Miika Aittala, Karsten Kreis +4

Diffusion models are a powerful tool for probabilistic forecasting, yet most applications in high-dimensional complex systems predict future states individually. This approach stru…

cs.LG2025

Improving Learning to Optimize Using Parameter Symmetries

Guy Zamir, Aryan Dokania, Bo Zhao +1

We analyze a learning-to-optimize (L2O) algorithm that exploits parameter space symmetry to enhance optimization efficiency. Prior work has shown that jointly learning symmetry tra…

cs.LG2019

Understanding the Representation Power of Graph Neural Networks in Learning Graph Topology

Nima Dehmamy, Albert-László Barabási, Rose Yu

To deepen our understanding of graph neural networks, we investigate the representation power of Graph Convolutional Networks (GCN) through the looking glass of graph moments, a ke…

cs.LG2018

Multi-resolution Tensor Learning for Large-Scale Spatial Data

Stephan Zheng, Rose Yu, Yisong Yue

High-dimensional tensor models are notoriously computationally expensive to train. We present a meta-learning algorithm, MMT, that can significantly speed up the process for spatia…

cs.LG2019

Long-term Forecasting using Higher Order Tensor RNNs

Rose Yu, Stephan Zheng, Anima Anandkumar +1

We present Higher-Order Tensor RNN (HOT-RNN), a novel family of neural sequence architectures for multivariate forecasting in environments with nonlinear dynamics. Long-term foreca…

cs.LG2025

Diffusion-BBO: Diffusion-Based Inverse Modeling for Online Black-Box Optimization

Dongxia Wu, Nikki Lijing Kuang, Ruijia Niu +2

Online black-box optimization (BBO) aims to optimize an objective function by iteratively querying a black-box oracle in a sample-efficient way. While prior studies focus on forwar…

cs.LG2026

U-Cast: A Surprisingly Simple and Efficient Frontier Probabilistic AI Weather Forecaster

Salva Rühling Cachay, Duncan Watson-Parris, Rose Yu

AI-based weather forecasting now rivals traditional physics-based ensembles, but state-of-the-art (SOTA) models rely on specialized architectures and massive computational budgets,…

cs.LG2025

Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation

Bohan Lyu, Yadi Cao, Duncan Watson-Parris +3

Large Language Models (LLMs) demonstrate promising capabilities in solving scientific problems but often suffer from the issue of hallucination. While integrating LLMs with tools c…

cs.LG2023

Symmetry Teleportation for Accelerated Optimization

Bo Zhao, Nima Dehmamy, Robin Walters +1

Existing gradient-based optimization methods update parameters locally, in a direction that minimizes the loss function. We study a different approach, symmetry teleportation, that…

cs.LG2022

Data Augmentation vs. Equivariant Networks: A Theory of Generalization on Dynamics Forecasting

Rui Wang, Robin Walters, Rose Yu

Exploiting symmetry in dynamical systems is a powerful way to improve the generalization of deep learning. The model learns to be invariant to transformation and hence is more robu…

cs.LG2021

Automatic Symmetry Discovery with Lie Algebra Convolutional Network

Nima Dehmamy, Robin Walters, Yanchen Liu +2

Existing equivariant neural networks require prior knowledge of the symmetry group and discretization for continuous groups. We propose to work with Lie algebras (infinitesimal gen…

cs.CL2026

Emergence of Hierarchical Emotion Organization in Large Language Models

Maya Okawa, Bo Zhao, Eric J. Bigelow +4

As large language models (LLMs) increasingly power conversational agents, understanding how they model users' emotional states is critical for ethical deployment. Inspired by emoti…

cs.LG2023

Physics-Guided Deep Learning for Dynamical Systems: A Survey

Rui Wang, Rose Yu

Modeling complex physical dynamics is a fundamental task in science and engineering. Traditional physics-based models are sample efficient, and interpretable but often rely on rigi…

cs.LG2026

Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success

Luca Zhou, Bo Zhao, Rose Yu +1

Model merging combines knowledge from separately fine-tuned models, yet the factors driving its success remain poorly understood. While recent work treats mergeability as an intrin…

cs.AI2026

Zephyrus: An Agentic Framework for Weather Science

Sumanth Varambally, Marshall Fisher, Jas Thakker +14

Foundation models for weather science are pre-trained on vast amounts of structured numerical data and outperform traditional weather forecasting systems. However, these models lac…

cs.LG2024

On the Theoretical Expressive Power and the Design Space of Higher-Order Graph Transformers

Cai Zhou, Rose Yu, Yusu Wang

Graph transformers have recently received significant attention in graph learning, partly due to their ability to capture more global interaction via self-attention. Nevertheless,…

cs.LG2023

On the Connection Between MPNN and Graph Transformer

Chen Cai, Truong Son Hy, Rose Yu +1

Graph Transformer (GT) recently has emerged as a new paradigm of graph learning algorithms, outperforming the previously popular Message Passing Neural Network (MPNN) on multiple b…

cs.LG2023

Automatic Integration for Spatiotemporal Neural Point Processes

Zihao Zhou, Rose Yu

Learning continuous-time point processes is essential to many discrete event forecasting tasks. However, integration poses a major challenge, particularly for spatiotemporal point…

cs.LG2018

Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting

Yaguang Li, Rose Yu, Cyrus Shahabi +1

Spatiotemporal forecasting has various applications in neuroscience, climate and transportation domain. Traffic forecasting is one canonical example of such learning task. The task…

cs.LG2026

Recursive Flow Matching

Jiahe Huang, Sihan Xu, Sharvaree Vadgama +1

Generative models have emerged as a powerful paradigm for solving physics systems and modeling complex spatiotemporal dynamics. However, achieving high physical accuracy without in…

cs.LG2024

ClimSim-Online: A Large Multi-scale Dataset and Framework for Hybrid ML-physics Climate Emulation

Sungduk Yu, Zeyuan Hu, Akshay Subramaniam +44

Modern climate projections lack adequate spatial and temporal resolution due to computational constraints, leading to inaccuracies in representing critical processes like thunderst…

cs.LG2023

Generative Adversarial Symmetry Discovery

Jianke Yang, Robin Walters, Nima Dehmamy +1

Despite the success of equivariant neural networks in scientific applications, they require knowing the symmetry group a priori. However, it may be difficult to know which symmetry…

cs.LG2022

LIMO: Latent Inceptionism for Targeted Molecule Generation

Peter Eckmann, Kunyang Sun, Bo Zhao +3

Generation of drug-like molecules with high binding affinity to target proteins remains a difficult and resource-intensive task in drug discovery. Existing approaches primarily emp…

cs.LG2023

DYffusion: A Dynamics-informed Diffusion Model for Spatiotemporal Forecasting

Salva Rühling Cachay, Bo Zhao, Hailey Joren +1

While diffusion models can successfully generate data and make predictions, they are predominantly designed for static images. We propose an approach for efficiently training diffu…

cs.LG2026

Discovering Symbolic Differential Equations with Symmetry Invariants

Jianke Yang, Manu Bhat, Bryan Hu +4

Discovering symbolic differential equations from data uncovers fundamental dynamical laws underlying complex systems. However, existing methods often struggle with the vast search…

cs.LG2022

Meta-Learning Dynamics Forecasting Using Task Inference

Rui Wang, Robin Walters, Rose Yu

Current deep learning models for dynamics forecasting struggle with generalization. They can only forecast in a specific domain and fail when applied to systems with different para…