3 papers
cs.LG2026
CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference
Xuezhen Xie, Zhiqiang Zhou
Large language model inference is bottlenecked by autoregressive decoding, where each token requires a full forward pass. Multi-token prediction (MTP) offers a promising accelerati…
cs.CV2026
Feature Alignment Determines Fusion Strategy: A Comparative Study of Cross-Attention and Concatenation in Multimodal Learning
Zhiqiang Zhou, Xuezhen Xie
The choice between cross-attention and concatenation for multimodal fusion remains governed by practitioner intuition rather than principled understanding. In this paper, we demons…
cs.CV2026
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
Enyi Shi, Pengyang Shao, Yanxin Zhang +5
The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current benchmarks typically isolate the…