collaborators

7 papers

cs.CV2025

MDIQA: Unified Image Quality Assessment for Multi-dimensional Evaluation and Restoration

Shunyu Yao, Ming Liu, Zhilu Zhang +4

Recent advancements in image quality assessment (IQA), driven by sophisticated deep neural network designs, have significantly improved the ability to approach human perceptions. H…

cs.AI2025

Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving

Zixian Guo, Ming Liu, Qilong Wang +4

Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end tra…

cs.CV2025

VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models

Kailai Feng, Yabo Zhang, Haodong Yu +4

Artistic typography is a technique to visualize the meaning of input character in an imaginable and readable manner. With powerful text-to-image diffusion models, existing methods…

cs.CG2025

SOLIDGEO: Measuring Multimodal Spatial Math Reasoning in Solid Geometry

Peijie Wang, Chao Yang, Zhong-Zhi Li +6

Geometry is a fundamental branch of mathematics and plays a crucial role in evaluating the reasoning capabilities of multimodal large language models (MLLMs). However, existing mul…

cs.CL2025

Enhancing Multimodal Continual Instruction Tuning with BranchLoRA

Duzhen Zhang, Yong Ren, Zhong-Zhi Li +5

Multimodal Continual Instruction Tuning (MCIT) aims to finetune Multimodal Large Language Models (MLLMs) to continually align with human intent across sequential tasks. Existing ap…

cs.CV2025

Explicit Relational Reasoning Network for Scene Text Detection

Yuchen Su, Zhineng Chen, Yongkun Du +4

Connected component (CC) is a proper text shape representation that aligns with human reading intuition. However, CC-based text detection methods have recently faced a developmenta…