activity
20242026
collaborators

9 papers

cs.RO2026

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Yihan Lin, Jiawei He, Shifeng Bao +6

Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAM…

cs.RO2026

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision

Haoyang Li, Guanlin Li, Youhe Feng +9

Cross-embodiment transfer in vision-language-action (VLA) models remains challenging because low-level state and action spaces differ fundamentally across robot platforms. We obser…

cs.SE2026

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

Jian Zhu, Yuzheng Zhang, Zeyao Ma +11

Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks evaluate isolated operations s…

cs.RO2026

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation

Youhe Feng, Hansen Shi, Haoyang Li +7

Long-horizon robotic manipulation requires dense feedback that reflects how a task advances through its procedural stages, not merely whether the final outcome is successful. Exist…

cs.CV2026

Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model

Chen Zhao, Zhuoran Wang, Haoyang Li +6

Vision-Language-Action (VLA) models have recently demonstrated strong performance across embodied tasks. Modern VLAs commonly employ diffusion action experts to efficiently generat…

cs.CL2025

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis

Bohan Zhang, Xiaokang Zhang, Jing Zhang +3

Current inference scaling methods, such as Self-consistency and Best-of-N, have proven effective in improving the accuracy of LLMs on complex reasoning tasks. However, these method…