activity
20242026
collaborators
Showing cs.LGShow all

13 papers · 1 filter

cs.LG2026

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation

Yu Fu, Longxuan Yu, Haz Sameen Shahgir +4

Safety alignment often improves robustness to harmful queries at the cost of reasoning ability, a tradeoff known as the safety tax. A common cause is distributional mismatch: super…

cs.LG2026

Sharpen Your Flow: Sharpness-Aware Sampling for Flow Matching

Aditi Gupta, Soon Hoe Lim, Annan Yu +1

Flow matching models generate samples by numerically integrating a learned velocity field, with each integration step requiring a neural network evaluation. Fast generation therefo…

cs.LG2026

Continuity Laws for Sequential Models

Annan Yu, Dongwei Lyu, N. Benjamin Erichson

Inductive biases influence the behavior and performance of sequential models. In this work, we study an underexplored inductive bias in sequential modeling: continuity in time. We…

cs.LG2026

Recency Biased Causal Attention for Time-series Forecasting

Kareem Hegazy, Michael W. Mahoney, N. Benjamin Erichson

Recency bias is a useful inductive prior for sequential modeling: it emphasizes nearby observations and can still allow longer-range dependencies. Standard Transformer attention la…

cs.LG2026

PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network Training

Shenghao Yang, Zhichao Wang, Oleg Balabanov +2

Matrix functions such as square root, inverse roots, and orthogonalization play a central role in preconditioned gradient methods for neural network training. This has motivated th…

cs.LG2025

WaveCastNet: Rapid Wavefield Forecasting for Earthquake Early Warning via Deep Sequence to Sequence Learning

Dongwei Lyu, Rie Nakata, Pu Ren +4

We propose a new deep learning model, WaveCastNet, to forecast high-dimensional wavefields. WaveCastNet integrates a convolutional long expressive memory architecture into a sequen…