collaborators

26 papers

cs.CV2026

LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection

Renshan Zhang, Haoyang Meng, Yixiao He +3

Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades sharply on small targets, de…

cs.CV2026

Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning

Yuxiang Xie, Qi Lv, Jianming Xing +4

Vision-language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpatialBrain, our submission to t…

cs.CV2026

Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction

Zijian Hong, Qi Lv, Yuxiang Xie +4

Visual pointing maps a language instruction to pixel co ordinates, a core skill for embodied AI. We describe our PointArena 2026 solution, which achieves 77.2% overall accuracy and…

cs.RO2026

Learning New Tasks via Reusable Skills: Skill-Compositional Experts for Embodied Continual Learning

Shuaike Zhang, Shaokun Wang, Haoyu Tang +2

Embodied Continual Learning (ECL) aims to enable robots to continually acquire new manipulation tasks while retaining previously learned behaviors under closed-loop control. Compar…

cs.CV2026

ViMax: Agentic Video Generation

Lingxuan Huang, Sizhe He, Hengji Zhou +3

Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Existing methods generate isolated sequence…

cs.CV2026

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

Wei Li, Renshan Zhang, Rui Shao +2

Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high computational overhead that limits…