activity
20232026
most citedMoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

4 citations · 8 across the 24 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models

Yijie Zhu, Zitong Yu, Wei Li +4

World Action Models (WAMs) extend Vision-Language-Action (VLA) models by incorporating future visual dynamics into action generation. However, existing WAMs often utilize imagined…

cs.CV2026

LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection

Renshan Zhang, Haoyang Meng, Yixiao He +3

Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades sharply on small targets, de…

cs.CV2026

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

Bing Hu, Zaijing Li, Rui Shao +4

Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varyi…

cs.CV2026

HATS: Hardness-Aware Trajectory Synthesis for GUI Agents

Rui Shao, Ruize Gao, Bin Xie +5

Graphical user interface (GUI) agents powered by large vision-language models (VLMs) have shown remarkable potential in automating digital tasks, highlighting the need for high-qua…

cs.CV2025

SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation

Wei Li, Renshan Zhang, Rui Shao +4

Vision-Language-Action (VLA) models have advanced in robotic manipulation, yet practical deployment remains hindered by two key limitations: 1) perceptual redundancy, where irrelev…

cs.CV2025

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

Wei Li, Renshan Zhang, Rui Shao +2

Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high computational overhead that limits…