978 citations · 1.3k across the 25 of their papers we have counts for
11 papers · 1 filter
PAPT++: Risk-Aware Adversarial Tuning and Generation for Single Domain Generalization
Zhipeng Xu, De Cheng, Xinyang Jiang +5
Single domain generalization (SDG) aims to learn a model from one labeled source domain that generalizes to unseen target domains. A common strategy is to enrich the source distrib…
ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval
Chunyi Peng, Zhipeng Xu, Yukun Yan +9
Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is…
DocClaw: A Unified Agentic System for Intelligent Document Processing
Siqi Xiang, Zhipeng Xu, Yufei Liu +6
Intelligent document processing (IDP) encompasses a broad range of tasks, including optical character recognition (OCR), document question answering (DocQA), and key information ex…
Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection
Mingyue Zeng, De Cheng, Zhipeng Xu +3
Incremental object detection (IOD) aims to extend detectors to new categories while retaining previously acquired knowledge. Existing methods often adopt a class incremental learni…
Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis
Zhipeng Xu, Zulong Chen, Qing Liu +6
Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-s…
Latent Visual Cache for Video Reasoning
Yongheng Zhang, Zhipeng Xu, Hao Wu +4
Video reasoning requires Large Multimodal Models (LMMs) to remain grounded in dense evidence, yet existing systems largely adopt "read-once, generate-many" paradigm, in which visua…