activity
20232026
most citedMoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

4 citations · 7 across the 29 of their papers we have counts for

collaborators
Showing cs.CVShow all

18 papers · 1 filter

cs.CV2026

Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models

Yijie Zhu, Zitong Yu, Wei Li +4

World Action Models (WAMs) extend Vision-Language-Action (VLA) models by incorporating future visual dynamics into action generation. However, existing WAMs often utilize imagined…

cs.CV2026

LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection

Renshan Zhang, Haoyang Meng, Yixiao He +3

Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades sharply on small targets, de…

cs.CV2026

ViMax: Agentic Video Generation

Lingxuan Huang, Sizhe He, Hengji Zhou +3

Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Existing methods generate isolated sequence…

cs.CV2026

Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning

Yuxiang Xie, Qi Lv, Jianming Xing +4

Vision-language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpatialBrain, our submission to t…

cs.CV2026

Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction

Zijian Hong, Qi Lv, Yuxiang Xie +4

Visual pointing maps a language instruction to pixel co ordinates, a core skill for embodied AI. We describe our PointArena 2026 solution, which achieves 77.2% overall accuracy and…

cs.CV2025

SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation

Wei Li, Renshan Zhang, Rui Shao +4

Vision-Language-Action (VLA) models have advanced in robotic manipulation, yet practical deployment remains hindered by two key limitations: 1) perceptual redundancy, where irrelev…