activity
20242026
most citedPrune4Web: DOM Tree Pruning Programming for Web Agent

1 citations · 1 across the 6 of their papers we have counts for

collaborators

13 papers

cs.RO2026

PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing

Yuheng Ji, Yuyang Liu, Huajie Tan +15

Current robotic evaluation is still largely dominated by binary success rates, which collapse rich execution processes into a single outcome and obscure critical qualities such as…

cs.CV2026

HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task

Xiaoya Lu, Yijin Zhou, Zeren Chen +6

Vision-Language Models (VLMs) empower embodied agents to execute complex instructions, yet they remain vulnerable to contextual safety risks where benign commands become hazardous…

cs.RO2026

SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics

Mengzhen Liu, Enshen Zhou, Cheng Chi +6

Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven active perception with robust, viewpoi…

cs.RO2026

RoboBrain 2.5: Depth in Sight, Time in Mind

Huajie Tan, Enshen Zhou, Zhiyu Li +32

We introduce RoboBrain 2.5, a next-generation embodied AI foundation model that advances general perception, spatial reasoning, and temporal modeling through extensive training on…

cs.CV2025

Towards Cross-View Point Correspondence in Vision-Language Models

Yipu Wang, Yuheng Ji, Yuyang Liu +10

Cross-view correspondence is a fundamental capability for spatial understanding and embodied AI. However, it is still far from being realized in Vision-Language Models (VLMs), espe…

cs.AI20251 cited

Prune4Web: DOM Tree Pruning Programming for Web Agent

Jiayuan Zhang, Kaiquan Chen, Zhihao Lu +3

Web automation employs intelligent agents to execute high-level tasks by mimicking human interactions with web interfaces. Despite the capabilities of recent Large Language Model (…