activity
20242026
collaborators

6 papers

cs.RO2026

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation

He Kong, Zengjue Chen, Qi Wang +6

Vision-language-action (VLA) models have demonstrated remarkable capabilities in robotic manipulation by leveraging pretrained vision-language models. However, existing post-traini…

cs.CV2026

From Bounding Boxes to Visual Reasoning: An On-Policy Data Annotation Tool for Vision-Language Models

Like Zhang, Runliang Niu, Shiqi Wang +5

Vision-language models (VLMs) are rapidly advancing toward sophisticated grounded structured visual reasoning. Training models for such advanced capabilities demands a new genre of…

cs.CV2026

Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision-Language Models

Yongguang Li, Jindong Li, Qi Wang +4

Vision-language models (VLMs) have gained widespread attention for their strong zero-shot capabilities across numerous downstream tasks. However, these models assume that each test…

cs.RO2025

TGRPO :Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy Optimization

Zengjue Chen, Runliang Niu, He Kong +3

Visual-Language-Action (VLA) models have demonstrated strong cross-scenario generalization capabilities in various robotic tasks through large-scale pre-training and task-specific…

cs.AI2025

ScreenExplorer: Training a Vision-Language Model for Diverse Exploration in Open GUI World

Runliang Niu, Jinglong Ji, Yi Chang +1

The rapid progress of large language models (LLMs) has sparked growing interest in building Artificial General Intelligence (AGI) within Graphical User Interface (GUI) environments…

cs.IR2024

Towards Next-Generation LLM-based Recommender Systems: A Survey and Beyond

Qi Wang, Jindong Li, Shiqi Wang +7

Large language models (LLMs) have not only revolutionized the field of natural language processing (NLP) but also have the potential to bring a paradigm shift in many other fields…