8 papers
VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
Yuqi Zhang, Cheng Chen, Yuyu Guo +6
Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions eit…
SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval
Wenjie Yang, Hang Yu, Yuyu Guo +1
The paper introduces SOLAR, a self‑supervised two‑stage framework for symmetric multimodal‑to‑multimodal retrieval that learns intersection masks from large unlabeled image‑text pa…
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
Shijie Li, Yilin Gao, Siyuan Yang +7
Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuanc…
N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization
Xukun Zhu, Hang Yu, Peng Di +1
The success of Large Language Models in mathematical reasoning relies heavily on the generation of diverse and valid solution paths during the rollout phase. However, current rollo…
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Ziyin Zhang, Zihan Liao, Hang Yu +2
The development of high-quality text embeddings is increasingly drifting toward an exclusionary future, defined by three critical barriers: prohibitive computational costs, a narro…
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
Xuexiang Wen, Hang Yu, Linchao Zhu +1
While Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a promising post-training paradigm for Large Language Models (LLMs), its dependency on the gold…