collaborators

14 papers

cs.CV2026

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Haojie Huang, Xinlei Yu, Chengming Xu +6

Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware pos…

cs.CV2026

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Jiahao Zhao, Xiaomin Yu, Zhongxiang Sun +5

Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, an…

cs.IR2026

SPEAR: Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval for Community Search

Wenbin Wu, Yuzhong Wu, Yufan Xu +4

Query reformulation bridges user intent and retrieval in e-commerce search, yet production systems optimize rewrite quality and retrieval effectiveness separately, leaving the two…

cs.AI2026

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

Kailin Lyu, Di Wu, Pengwei Zhang +12

Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense re…

cs.RO2026

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6

Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally de…

cs.CV2026

SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs

Bo Yin, Xiaobin Hu, Chengming Xu +6

Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evi…