collaborators

7 papers

cs.CL2026

Stopping and Routing LLM Judge Panels

Bin Zhu, Yi Xie, Yanghui Rao

LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The…

cs.CV2026

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

Huiqiong Li, Jiayu Wang, Zhiting Mei +3

Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instructions. We introduce RoboTrustB…

cs.CL2026

A Finite-Calibration Regime Map for LLM Judge Panels

Bin Zhu, Yi Xie, Yanghui Rao

Deploying an LLM judge panel spends human labels on fitting a calibrator, constructing candidate judge paths, and validating which candidate to deploy. We study when finite labels…

cs.RO2026

-WM: A Unified Video-Action World Model for Robotic Manipulation

Pengfei Zhou, Shengcong Chen, Di Chen +17

Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present -World…

cs.RO2026

Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA

Hung Mai, Bin Zhu, Tuan Do

Vision-language-action (VLA) policies and World-Action Models (WAM) represent two increasingly important paradigms for robotic manipulation. However, it remains unclear whether fut…

cs.CV2026

Teacher-Student Diffusion Model for Text-Driven 3D Hand Motion Generation

Ching-Lam Cheng, Bin Zhu, Shengfeng He

Generating realistic 3D hand motion from natural language is vital for VR, robotics, and human-computer interaction. Existing methods either focus on full-body motion, overlooking…