From the 1 of 7 linked papers with an AI index.
7 papers
SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling
Zheng Liu, Zijian He, Huiguo He +5
Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span differ…
JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis
Ran Li, Huiguo He, Jiahuan Cao +3
The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches…
TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents
Zhongheng Zhou, Yi Sun, Huiguo He +6
Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and…
One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting
Rui Tang, Wentao Yang, Peirong Zhang +4
The paper introduces a vision-centric framework called SPaTS that uses a single visual token per text instance and reinforcement learning to improve scene text spotting accuracy an…
DiffMath: Symbol- and Graph-Aware Latent Diffusion Transformer for Handwritten Mathematical Expression Generation
Wei Pan, Xuhan Zheng, Yilin Shi +5
Handwritten Mathematical Expression Generation (HMEG) is challenging due to the complex two-dimensional layouts and long-range structural dependencies of mathematical expressions.…
ContextDrag: Precise Drag-Based Image Editing via Context-Preserving Token Injection and Position-Aligned Attention
Huiguo He, Pengyu Yan, Ziqi Yi +6
Drag-based image editing enables intuitive visual manipulation through point-based drag operations. Existing methods mainly rely on diffusion inversion or pixel-space warping with…