collaborators

5 papers

cs.CV2026

OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios

Zinuo Guo, Min Zhang, Bo Jiang

Multimodal large language models (MLLMs) are increasingly used as OCR systems in document and knowledge-processing pipelines, but their ability to faithfully read real handwriting…

cs.CV2026

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models

Xinpeng Dong, Min Zhang, Kairong Han +3

In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual informat…

cs.MA2026

MetaForge: A Self-Evolving Multimodal Agent that Retrieves, Adapts, and Forges Tools On Demand

Shouang Wei, Houcheng Min, Xinpeng Dong +8

Multimodal agents have achieved notable progress on complex reasoning tasks through tool use, yet remain limited by two issues: statically predefined tool inventories fail to gener…

cs.CL2026

CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models

Rui Jia, Ruiyi Lan, Fengrui Liu +7

Large language models (LLMs) have advanced the development of personalized learning in education. However, their inherent generation mechanisms often produce homogeneous responses…

cs.CL2026

Reversible Diffusion Decoding for Diffusion Language Models

Xinyun Wang, Min Zhang, Sen Cui +4

Diffusion language models enable parallel token generation through block-wise decoding, but their irreversible commitments can lead to stagnation, where the reverse diffusion proce…