7 papers
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
Zeyu Chen, Huanjin Yao, Ziwang Zhao +1
The paper introduces a new benchmark, M-JudgeBench, to evaluate the judgment capabilities of multimodal large language models, and proposes a data generation method (Judge-MCTS) to…
H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation
Qixiang Yin, Huanjin Yao, Yuchen Cai +5
On-policy distillation (OPD) has recently emerged as an effective post-training paradigm by providing supervision on student-generated trajectories. However, existing OPD methods f…
Valley3: Scaling Omni Foundation Models for E-commerce
Zeyu Chen, Guanghao Zhou, Qixiang Yin +6
In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understanding and reasoning capabilitie…
Improving Vision-language Models with Perception-centric Process Reward Models
Yingqian Min, Kun Zhou, Yifan Li +6
Recent advancements in reinforcement learning with verifiable rewards (RLVR) have significantly improved the complex reasoning ability of vision-language models (VLMs). However, it…
MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline
Huanjin Yao, Qixiang Yin, Min Yang +5
We aim to develop a multimodal research agent capable of explicit reasoning and planning, multi-tool invocation, and cross-modal information synthesis, enabling it to conduct deep…
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
Yifan Li, Yukai Gu, Yingqian Min +6
Recent breakthroughs in video generation have demonstrated an emerging capability termed Chain-of-Frames (CoF) reasoning, where models resolve complex tasks through the generation…