collaborators

25 papers

cs.CV2026

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Jiahao Zhao, Xiaomin Yu, Zhongxiang Sun +5

Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, an…

cs.AI2026

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

Kailin Lyu, Di Wu, Pengwei Zhang +12

Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense re…

cs.CV2026

FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning

Bo Yin, Xiaobin Hu, Xingyu Zhou +7

Diffusion models have achieved remarkable success in generative modeling, yet how to effectively adapt large pretrained models to new tasks remains challenging. We revisit the reco…

eess.IV2026

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

Juntao Jiang, Jiangning Zhang, Yali Bi +7

Chain-of-Thought (CoT) reasoning has proven effective in enhancing large language models by encouraging step-by-step intermediate reasoning, and recent advances have extended this…

cs.CV2026

MambaADv2: Evolving Duality-enhanced State Space Model for Unsupervised Anomaly Detection

Xiaobin Hu, Haoyang He, Bo Yin +5

While recent advancements in anomaly detection have demonstrated the efficacy of CNN- and Transformer-based approaches, these architectures face inherent limitations: CNNs struggle…

cs.CV2026

SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs

Bo Yin, Xiaobin Hu, Chengming Xu +6

Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evi…