collaborators

5 papers

cs.RO2026

OrthoSkillVLA: Continual Skill Learning via Gradient-Informed Skill Subspace Adaptation

Jiaqi Wang, Zhou Fang, Qiongfeng Shi +1

Pretrained Vision-Language-Action models provide a strong foundation for robot learning, but sequentially adapting them to diverse skills can perturb the representations and veloci…

cs.AI2026

Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach

Ruichao Mao, Zhou Fang, Teng Guo +9

User experience (UX) centered on usability, perceived consistency, and functional clarity is fundamental to real-world user interfaces (UI). The application of multimodal large lan…

cs.AI2026

Towards Scalable Lightweight GUI Agents via Multi-role Orchestration

Ziwei Wang, Junjie Zheng, Leyang Yang +7

Autonomous Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) enable digital automation on end-user devices. While scaling both parameters an…

cs.RO2026

ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models

Zhou Fang, Jiaqi Wang, Yi Zhou +1

Recent Vision-Language-Action (VLA) models equipped with Flow Matching (FM) action heads achieve state-of-the-art performance in complex robot manipulation. However, the multi-step…

cs.LG2024

vMFER: Von Mises-Fisher Experience Resampling Based on Uncertainty of Gradient Directions for Policy Improvement

Yiwen Zhu, Jinyi Liu, Wenya Wei +7

Reinforcement Learning (RL) is a widely employed technique in decision-making problems, encompassing two fundamental operations -- policy evaluation and policy improvement. Enhanci…