activity
20172026
most citedMMOCR: A Comprehensive Toolbox for Text Detection, Recognition and Understanding

15 citations · 33 across the 7 of their papers we have counts for

collaborators

10 papers

cs.CV2026

Empirical Recipes for Efficient and Compact Vision-Language Models

Jiabo Huang, Zhizhong Li, Sina Sajadmanesh +2

Deploying vision-language models (VLMs) in resource-constrained settings demands low latency and high throughput, yet existing compact VLMs often fall short of the inference speedu…

cs.LG2025

StelLA: Subspace Learning in Low-rank Adaptation using Stiefel Manifold

Zhizhong Li, Sina Sajadmanesh, Jingtao Li +1

Low-rank adaptation (LoRA) has been widely adopted as a parameter-efficient technique for fine-tuning large-scale pre-trained models. However, it still lags behind full fine-tuning…

cs.CV2025

ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model

Weitai Kang, Weiming Zhuang, Zhizhong Li +2

Fine-grained multimodal capability in Multimodal Large Language Models (MLLMs) has emerged as a critical research direction, particularly for tackling the visual grounding (VG) pro…

cs.CV202115 cited

MMOCR: A Comprehensive Toolbox for Text Detection, Recognition and Understanding

Zhanghui Kuang, Hongbin Sun, Zhizhong Li +10

We present MMOCR-an open-source toolbox which provides a comprehensive pipeline for text detection and recognition, as well as their downstream tasks such as named entity recogniti…

cs.CV20212 cited

Representation Consolidation for Training Expert Students

Zhizhong Li, Avinash Ravichandran, Charless Fowlkes +3

Traditionally, distillation has been used to train a student model to emulate the input/output functionality of a teacher. A more useful goal than emulation, yet under-explored, is…

cs.CV20208 cited

Regularizing Reasons for Outfit Evaluation with Gradient Penalty

Xingxing Zou, Zhizhong Li, Ke Bai +2

In this paper, we build an outfit evaluation system which provides feedbacks consisting of a judgment with a convincing explanation. The system is trained in a supervised manner wh…