collaborators

6 papers

cs.LG2026

BRIDGE: Bridging Reasoning In Distillation Gap Elimination via Structure-Aware Masking

Bowen Yu, Sheng Zhang, Binhao Wang +8

Chain-of-Thought (CoT) reasoning has significantly improved LLMs' mathematical problem-solving capabilities, but distilling such capabilities into smaller models remains challengin…

cs.CL2026

CSCBench: A PVC Diagnostic Benchmark for Commodity Supply Chain Reasoning

Yaxin Cui, Yuanqiang Zeng, Jiapeng Yan +8

Large Language Models (LLMs) have achieved remarkable success in general benchmarks, yet their competence in commodity supply chains (CSCs) -- a domain governed by institutional ru…

cs.RO2025

A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning

Shaopeng Zhai, Qi Zhang, Tianyi Zhang +7

Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration. We introduce VLA…

cs.RO2025

Embodied Long Horizon Manipulation with Closed-loop Code Generation and Incremental Few-shot Adaptation

Yuan Meng, Xiangtong Yao, Haihui Ye +6

Embodied long-horizon manipulation requires robotic systems to process multimodal inputs-such as vision and natural language-and translate them into executable actions. However, ex…

cs.LG2025

Test-Time Scaling with Reflective Generative Model

Zixiao Wang, Yuxin Wang, Xiaorui Wang +8

We introduce our first reflective generative model MetaStone-S1, which obtains OpenAI o3-mini's performance via the new Reflective Generative Form. The new form focuses on high-qua…

cs.RO2025

From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control

Jusheng Zhang, Jinzhou Tang, Sidi Liu +4

Human motion generative modeling or synthesis aims to characterize complicated human motions of daily activities in diverse real-world environments. However, current research predo…