5 papers
Learning Variable-Length Tokenization for Generative Recommendation
Minhao Wang, Bowen Wu, Wei Zhang
Generative recommendation reformulates recommendation as next-token prediction over discrete semantic identifiers (IDs). A fundamental yet unexplored design choice is that existing…
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution
Zehao Wang, Yihan Zeng, Zidong Gong +5
Post-training via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) is crucial for enhancing reasoning in Multimodal Large Language Models (MLLMs), yet existing paradigm…
Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey
Liangwei Nathan Zheng, Wei Emma Zhang, Olaf Maennel +2
Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across diverse modalities and tasks. Desp…
Surface-Form Neural Sparse Retrieval: Robust Fuzzy Matching for Industrial Music Search
Paul Greyson, Zhichao Geng, Wei Zhang +1
Music search at the scale of Amazon Music presents a unique challenge: queries frequently deviate from indexed metadata due to misspellings, transpositions, and phonetic variations…
MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
Peng Xu, Shengwu Xiong, Jiajun Zhang +125
This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…