4 citations · 4 across the 16 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation
Jingang Zhou, Yuyi Zhou, Haiyang Guo +6
On-policy distillation (OPD) is a promising approach for transferring knowledge between language models, where a student receives dense token-level supervision along its own genera…
cs.LG2026
Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models
Ruihan Xu, Yuting Gao, Lan Wang +5
Large Multimodal Models (LMMs) have achieved remarkable success in vision-language tasks, yet their vast parameter counts are often underutilized during both training and inference…
cs.LG2025
M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance
Qingpei Guo, Kaiyou Song, Zipeng Feng +9
We present M2-omni, a cutting-edge, open-source omni-MLLM that achieves competitive performance to GPT-4o. M2-omni employs a unified multimodal sequence modeling framework, which e…