13 papers · 1 filter
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
Yu Fu, Longxuan Yu, Haz Sameen Shahgir +4
Safety alignment often improves robustness to harmful queries at the cost of reasoning ability, a tradeoff known as the safety tax. A common cause is distributional mismatch: super…
Sharpen Your Flow: Sharpness-Aware Sampling for Flow Matching
Aditi Gupta, Soon Hoe Lim, Annan Yu +1
Flow matching models generate samples by numerically integrating a learned velocity field, with each integration step requiring a neural network evaluation. Fast generation therefo…
Continuity Laws for Sequential Models
Annan Yu, Dongwei Lyu, N. Benjamin Erichson
Inductive biases influence the behavior and performance of sequential models. In this work, we study an underexplored inductive bias in sequential modeling: continuity in time. We…
Recency Biased Causal Attention for Time-series Forecasting
Kareem Hegazy, Michael W. Mahoney, N. Benjamin Erichson
Recency bias is a useful inductive prior for sequential modeling: it emphasizes nearby observations and can still allow longer-range dependencies. Standard Transformer attention la…
PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network Training
Shenghao Yang, Zhichao Wang, Oleg Balabanov +2
Matrix functions such as square root, inverse roots, and orthogonalization play a central role in preconditioned gradient methods for neural network training. This has motivated th…
WaveCastNet: Rapid Wavefield Forecasting for Earthquake Early Warning via Deep Sequence to Sequence Learning
Dongwei Lyu, Rie Nakata, Pu Ren +4
We propose a new deep learning model, WaveCastNet, to forecast high-dimensional wavefields. WaveCastNet integrates a convolutional long expressive memory architecture into a sequen…