6 papers
Learning Variable-Length Tokenization for Generative Recommendation
Minhao Wang, Bowen Wu, Wei Zhang
Generative recommendation reformulates recommendation as next-token prediction over discrete semantic identifiers (IDs). A fundamental yet unexplored design choice is that existing…
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution
Zehao Wang, Yihan Zeng, Zidong Gong +5
Post-training via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) is crucial for enhancing reasoning in Multimodal Large Language Models (MLLMs), yet existing paradigm…
World Models as Group Actions
Zijie Wang, Wei Zhang, Weiming Zhang +4
Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness…
Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey
Liangwei Nathan Zheng, Wei Emma Zhang, Olaf Maennel +2
Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across diverse modalities and tasks. Desp…
Surface-Form Neural Sparse Retrieval: Robust Fuzzy Matching for Industrial Music Search
Paul Greyson, Zhichao Geng, Wei Zhang +1
Music search at the scale of Amazon Music presents a unique challenge: queries frequently deviate from indexed metadata due to misspellings, transpositions, and phonetic variations…
MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
Peng Xu, Shengwu Xiong, Jiajun Zhang +125
This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…