3 papers
cs.CV2026
Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction
Yichang Jian, Boyuan Xiao, Zhenyuan Huang +2
Planning from raw visual input remains a significant challenge for current Vision-Language Models (VLMs), when the complexity of input is beyond their one-step perception capabilit…
cs.IR2023
Workshop on Document Intelligence Understanding
Soyeon Caren Han, Yihao Ding, Siwen Luo +5
Document understanding and information extraction include different tasks to understand a document and extract valuable information automatically. Recently, there has been a rising…
cs.AI2022
V-Doc : Visual questions answers with Documents
Yihao Ding, Zhe Huang, Runlin Wang +5
We propose V-Doc, a question-answering tool using document images and PDF, mainly for researchers and general non-deep learning experts looking to generate, process, and understand…