21 papers
LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection
Renshan Zhang, Haoyang Meng, Yixiao He +3
Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades sharply on small targets, de…
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
Yibo Lyu, Rui Shao, Gongwei Chen +3
As multimedia content expands, the demand for unified multimodal retrieval (UMR) in real-world applications increases. Recent work leverages multimodal large language models (MLLMs…
From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model
Bing Hu, Zaijing Li, Rui Shao +4
Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varyi…
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
Wei Li, Renshan Zhang, Rui Shao +2
Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high computational overhead that limits…
UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
Yijie Zhu, Lingsen Zhang, Zitong Yu +3
Emotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each other. In this paper, we propose the…
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
Zaijing Li, Bing Hu, Rui Shao +5
Hierarchical Vision-Language-Action (VLA) models have rapidly become a dominant paradigm for robotic manipulation. It typically comprising a Vision-Language backbone for perception…