activity
20232026
most citedSteve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds

4 citations · 5 across the 13 of their papers we have counts for

collaborators
Showing 2026Show all

6 papers · 1 filter

cs.AI2026

Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents

AIMAE Team, Tianxiang Chen, Yan Cheng +39

Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of decisions on live websites while recoverin…

cs.RO2026

Being-H0.7: A Latent World-Action Model from Egocentric Videos

Hao Luo, Wanpeng Zhang, Yicheng Feng +6

Visual-Language-Action models (VLAs) have advanced generalist robot control by mapping multimodal observations and language instructions directly to actions, but sparse action supe…

cs.RO2026

Conservative Offline Robot Policy Learning via Posterior-Transition Reweighting

Wanpeng Zhang, Hao Luo, Sipeng Zheng +6

Offline post-training adapts a pretrained robot policy to a target dataset by supervised regression on recorded actions. In practice, robot datasets are heterogeneous: they mix emb…

cs.RO2026

Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild

Hao Luo, Ye Wang, Wanpeng Zhang +5

Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, exist…

cs.RO2026

Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization

Ye Wang, Sipeng Zheng, Hao Luo +9

While Vision-Language-Action (VLA) models show strong promise for generalist robot control, it remains unclear whether -- and under what conditions -- the standard "scale data" rec…

cs.RO2026

Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization

Hao Luo, Ye Wang, Wanpeng Zhang +9

We introduce Being-H0.5, a foundational Vision-Language-Action (VLA) model designed for robust cross-embodiment generalization across diverse robotic platforms. While existing VLAs…