8 papers
DualMem: Bypassing the Objectness Bottleneck for Calibrated Unknown-Stream Filtering in Open-World Object Detection
Yingjun Xiao, Xi Chen, Gang Fang +1
Open-world object detection (OWOD) requires detectors to localize known classes while identifying unknown objects for future incremental learning. We find that the unknown predicti…
Attention Sinks and Outliers in Attention Residuals
Haozheng Luo, Haoran Dai, Shaoyang Zhang +10
We propose OASIS, an outlier- and sink-aware technique built on inter-layer null signaling. As AttnResidual architectures introduce an additional depth-wise normalization channel,…
Evidence-Guided Unknown Rejection for High-Confidence Near-Known Unknowns
Xi Chen, Yingjun Xiao, Gang Fang
Open-set recognition systems face a neglected failure mode: high-confidence near-known unknowns, which lie outside the known label set but are close enough to known classes that a…
Optimal low-rank stochastic gradient estimation for LLM training
Zehao Li, Tao Ren, Zishi Zhang +2
Large language model (LLM) training is often bottlenecked by memory constraints and stochastic gradient noise in extremely high-dimensional parameter spaces. Motivated by empirical…
Astro: Activation-guided Structured Regularization for Outlier-Robust LLM Post-Training Quantization
Xi Chen, Ming Li, Junxi Li +5
Weight-only post-training quantization (PTQ) is crucial for efficient Large Language Model (LLM) deployment but suffers from accuracy degradation caused by weight and activation ou…
Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural Networks
Shenxi Wu, Haosong Zhang, Xingjian Ma +4
Deeper modern architectures are costly to train, making hyperparameter transfer preferable to expensive repeated tuning. Maximal Update Parametrization (P) helps explain why ma…