activity
20212026
most citedRead Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition

21 citations · 36 across the 5 of their papers we have counts for

collaborators

8 papers

cs.CV2026

Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement

Junrong Guo, Shancheng Fang, Yadong Qu +1

Recent advances in Multimodal Large Language Models (MLLMs) have enabled automated generation of structured layouts from natural language descriptions. Existing methods typically f…

cs.IR2026

FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with Citations

Yixing Peng, Licheng Zhang, Shancheng Fang +3

Generating with citations is crucial for trustworthy Large Language Models (LLMs), yet even advanced LLMs often produce mismatched or irrelevant citations. Existing methods over-op…

cs.CV2025

IGD: Instructional Graphic Design with Multimodal Layer Generation

Yadong Qu, Shancheng Fang, Yuxin Wang +4

Graphic design visually conveys information and data by creating and combining text, images and graphics. Two-stage methods that rely primarily on layout generation lack creativity…

cs.CV2025

MaskDiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

Tianhao Qi, Jianlong Yuan, Wanquan Feng +6

Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video ge…

cs.CL2024

GRIP: A Graph-Based Reasoning Instruction Producer

Jiankang Wang, Jianjun Xu, Xiaorui Wang +4

Large-scale, high-quality data is essential for advancing the reasoning capabilities of large language models (LLMs). As publicly available Internet data becomes increasingly scarc…

cs.CV20219 cited

From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network

Yuxin Wang, Hongtao Xie, Shancheng Fang +3

In this paper, we abandon the dominant complex language model and rethink the linguistic learning process in the scene text recognition. Different from previous methods considering…