From the 1 of 18 linked papers with an AI index.
18 papers
SceneActBench: Can Agents Act on the 3D Scenes They See?
Yifei Zhao, Xiangxin Zhou, Wenhao Yang +11
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operat…
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
Zeyu Chen, Huanjin Yao, Ziwang Zhao +1
The paper introduces a new benchmark, M-JudgeBench, to evaluate the judgment capabilities of multimodal large language models, and proposes a data generation method (Judge-MCTS) to…
H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation
Qixiang Yin, Huanjin Yao, Yuchen Cai +5
On-policy distillation (OPD) has recently emerged as an effective post-training paradigm by providing supervision on student-generated trajectories. However, existing OPD methods f…
OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models
Tianyi Lin, Chuanyu Sun, Jingyi Zhang +6
Equipping Large Language Model (LLM) agents with effective skills is crucial for solving complex tasks in real-world systems like OpenClaw. In this work, we aim to develop a framew…
R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?
Jingyi Zhang, Tianyi Lin, Huanjin Yao +3
In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks. T…
Valley3: Scaling Omni Foundation Models for E-commerce
Zeyu Chen, Guanghao Zhou, Qixiang Yin +6
In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understanding and reasoning capabilitie…