collaborators

6 papers

cs.RO2026

ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning

Wei Xiao, Weiliang Tang, Yuying Ge +4

Human interventions provide crucial corrective signals for post-training Vision-Language-Action (VLA) models. However, enabling seamless humanoid interventions is a formidable syst…

cs.CL2026

PathRouter: Aligning Rewards with Retrieval Quality in Agentic Graph Retrieval-Augmented Generation

Bo Wang, Heyan Huang, Yaolin Li +6

Agentic GraphRAG trains language-model agents to iteratively retrieve and reason over graph-structured evidence, enabling more accurate and context-aware decision-making by efficie…

cs.CV2026

Artemis: Structured Visual Reasoning for Perception Policy Learning

Wei Tang, Yanpeng Sun, Shan Zhang +6

Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural language. Empirical observations indica…

cs.GR2026

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

Lin Song, Wenbo Li, Guoqing Ma +16

We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatia…

cs.CV2026

SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation

Wei Tang, Xuejing Liu, Yanpeng Sun +1

The Segment Anything Model (SAM) excels at general image segmentation but has limited ability to understand natural language, which restricts its direct application in Referring Ex…

cs.CV2025

Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs

Yanpeng Sun, Shan Zhang, Wei Tang +5

Diagrams represent a form of visual language that encodes abstract concepts and relationships through structured symbols and their spatial arrangements. Unlike natural images, they…