4 citations · 7 across the 29 of their papers we have counts for
18 papers · 1 filter
Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models
Yijie Zhu, Zitong Yu, Wei Li +4
World Action Models (WAMs) extend Vision-Language-Action (VLA) models by incorporating future visual dynamics into action generation. However, existing WAMs often utilize imagined…
LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection
Renshan Zhang, Haoyang Meng, Yixiao He +3
Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades sharply on small targets, de…
ViMax: Agentic Video Generation
Lingxuan Huang, Sizhe He, Hengji Zhou +3
Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Existing methods generate isolated sequence…
Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning
Yuxiang Xie, Qi Lv, Jianming Xing +4
Vision-language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpatialBrain, our submission to t…
Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction
Zijian Hong, Qi Lv, Yuxiang Xie +4
Visual pointing maps a language instruction to pixel co ordinates, a core skill for embodied AI. We describe our PointArena 2026 solution, which achieves 77.2% overall accuracy and…
SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
Wei Li, Renshan Zhang, Rui Shao +4
Vision-Language-Action (VLA) models have advanced in robotic manipulation, yet practical deployment remains hindered by two key limitations: 1) perceptual redundancy, where irrelev…