26 papers
LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection
Renshan Zhang, Haoyang Meng, Yixiao He +3
Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades sharply on small targets, de…
Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning
Yuxiang Xie, Qi Lv, Jianming Xing +4
Vision-language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpatialBrain, our submission to t…
Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction
Zijian Hong, Qi Lv, Yuxiang Xie +4
Visual pointing maps a language instruction to pixel co ordinates, a core skill for embodied AI. We describe our PointArena 2026 solution, which achieves 77.2% overall accuracy and…
Learning New Tasks via Reusable Skills: Skill-Compositional Experts for Embodied Continual Learning
Shuaike Zhang, Shaokun Wang, Haoyu Tang +2
Embodied Continual Learning (ECL) aims to enable robots to continually acquire new manipulation tasks while retaining previously learned behaviors under closed-loop control. Compar…
ViMax: Agentic Video Generation
Lingxuan Huang, Sizhe He, Hengji Zhou +3
Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Existing methods generate isolated sequence…
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
Wei Li, Renshan Zhang, Rui Shao +2
Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high computational overhead that limits…