papers

Publications (6)

cs.CV2022

UFO: Unified Feature Optimization

Teng Xi, Yifan Sun, Deli Yu +13

This paper proposes a novel Unified Feature Optimization (UFO) paradigm for training and deploying deep models under real-world and large-scale scenarios, which requires a collecti…

cs.CV2025

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration

Fanhu Zeng, Deli Yu, Zhenglun Kong +1

Vision transformers have been widely explored in various vision tasks. Due to heavy computational cost, much interest has aroused for compressing vision transformer dynamically in…

cs.CL2026

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs

Baorong Shi, Bo Cui, Boyuan Jiang +17

We present MedXIAOHE, a medical vision-language foundation model designed to advance general-purpose medical understanding and reasoning in real-world clinical applications. MedXIA…

cs.CV2020

Towards Accurate Scene Text Recognition with Semantic Reasoning Networks

Deli Yu, Xuan Li, Chengquan Zhang +3

Scene text image contains two levels of contents: visual texture and semantic information. Although the previous scene text recognition methods have made great progress over the pa…

cs.CV2023

Accelerating Vision Transformers Based on Heterogeneous Attention Patterns

Deli Yu, Teng Xi, Jianwei Li +7

Recently, Vision Transformers (ViTs) have attracted a lot of attention in the field of computer vision. Generally, the powerful representative capacity of ViTs mainly benefits from…

cs.CL2026

Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs

Jianghang Lin, Haihua Yang, Deli Yu +6

Multimodal Large Language Models (MLLMs) have shown transformative potential in medical applications, yet their performance is hindered by conventional data curation strategies tha…