3 papers
cs.CL2026
Depth-Attention: Cross-Layer Value Mixing for Language Models
Boyi Zeng, Yiqin Hao, Zitong Wang +7
Self-attention selects information freely across the sequence, but across depth, Transformers merely add each layer's output to the residual stream, so later layers cannot selectiv…
cs.CV2026
Label What Matters: Modality-Balanced and Difficulty-Aware Multimodal Active Learning
Yuqiao Zeng, Xu Wang, Tengfei Liang +3
Multimodal learning integrates complementary information from different modalities such as image, text, and audio to improve model performance, but its success relies on large-scal…
cs.CL2026
Pretraining with Token-Level Adaptive Latent Chain-of-Thought
Boyi Zeng, Yiqin Hao, He Li +8
Scaling large language models by increasing parameters and training data is increasingly constrained by limited high-quality corpora and rising communication costs. This work explo…