2 papers
cs.CV2026
DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding
Jianfei Zhao, Feng Zhang, Xin Sun +3
Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to text tokens. Visual tokens, th…
cs.LG2026
Attention Sinks and Outliers in Attention Residuals
Haozheng Luo, Haoran Dai, Shaoyang Zhang +10
We propose OASIS, an outlier- and sink-aware technique built on inter-layer null signaling. As AttnResidual architectures introduce an additional depth-wise normalization channel,…