collaborators

6 papers

cs.LG2026

Learning Variable-Length Tokenization for Generative Recommendation

Minhao Wang, Bowen Wu, Wei Zhang

Generative recommendation reformulates recommendation as next-token prediction over discrete semantic identifiers (IDs). A fundamental yet unexplored design choice is that existing…

cs.CV2026

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution

Zehao Wang, Yihan Zeng, Zidong Gong +5

Post-training via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) is crucial for enhancing reasoning in Multimodal Large Language Models (MLLMs), yet existing paradigm…

cs.CV2026

World Models as Group Actions

Zijie Wang, Wei Zhang, Weiming Zhang +4

Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness…

cs.LG2026

Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey

Liangwei Nathan Zheng, Wei Emma Zhang, Olaf Maennel +2

Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across diverse modalities and tasks. Desp…

cs.AI2026

Surface-Form Neural Sparse Retrieval: Robust Fuzzy Matching for Industrial Music Search

Paul Greyson, Zhichao Geng, Wei Zhang +1

Music search at the scale of Amazon Music presents a unique challenge: queries frequently deviate from indexed metadata due to misspellings, transpositions, and phonetic variations…

cs.CV2025

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

Peng Xu, Shengwu Xiong, Jiajun Zhang +125

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…