collaborators

23 papers

cs.CV2026

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

Runhui Huang, Qihui Zhang, Zhe Liu +3

The paper introduces SpectraReward, a training-free method that uses pretrained multimodal large language models to score generated images by measuring how well the original text p…

cs.RO2026

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI

Brain Team, Ziyang Gong, Haoming Gu +28

Embodied AI is moving from isolated perception or action modules toward physical agents that understand, plan under goals, act through robot bodies, monitor progress, and improve f…

cs.AI2026

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting

Junwei Luo, Shuai Yuan, Zhenya Yang +3

Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. In this paper, we view this t…

cs.RO2026

Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention

Siyu Wu, Linjing You, Junjie Zhu +10

World Action Models (WAMs) jointly predict future visual observations and actions, but visual futures alone often miss slip, jamming, contact-direction changes, and subtle misalign…

cs.AI2026

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning

Xirui Li, Zhe Liu, Xiaoqing Ye +4

Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervision but are confined to a fixed action voca…

cs.CV2026

SCOPE: Scale-Consistent One-Pass Estimation of 3D Geometry

Zheng Zhang, Lihe Yang, Tianyu Yang +6

We present SCOPE (Scale-Consistent One-Pass Estimation of 3D Geometry), a novel approach for estimating 3D geometry from extended monocular video sequences, where existing methods…