collaborators

21 papers

cs.CV2026

LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection

Renshan Zhang, Haoyang Meng, Yixiao He +3

Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades sharply on small targets, de…

cs.MM2026

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning

Yibo Lyu, Rui Shao, Gongwei Chen +3

As multimedia content expands, the demand for unified multimodal retrieval (UMR) in real-world applications increases. Recent work leverages multimodal large language models (MLLMs…

cs.CV2026

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

Bing Hu, Zaijing Li, Rui Shao +4

Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varyi…

cs.CV2026

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

Wei Li, Renshan Zhang, Rui Shao +2

Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high computational overhead that limits…

cs.CV2026

UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries

Yijie Zhu, Lingsen Zhang, Zitong Yu +3

Emotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each other. In this paper, we propose the…

cs.RO2026

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Zaijing Li, Bing Hu, Rui Shao +5

Hierarchical Vision-Language-Action (VLA) models have rapidly become a dominant paradigm for robotic manipulation. It typically comprising a Vision-Language backbone for perception…