activity
20242026
most citedDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

978 citations · 1.3k across the 25 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

PAPT++: Risk-Aware Adversarial Tuning and Generation for Single Domain Generalization

Zhipeng Xu, De Cheng, Xinyang Jiang +5

Single domain generalization (SDG) aims to learn a model from one labeled source domain that generalizes to unseen target domains. A common strategy is to enrich the source distrib…

cs.CV2026

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

Chunyi Peng, Zhipeng Xu, Yukun Yan +9

Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is…

cs.CV2026

DocClaw: A Unified Agentic System for Intelligent Document Processing

Siqi Xiang, Zhipeng Xu, Yufei Liu +6

Intelligent document processing (IDP) encompasses a broad range of tasks, including optical character recognition (OCR), document question answering (DocQA), and key information ex…

cs.CV2026

Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection

Mingyue Zeng, De Cheng, Zhipeng Xu +3

Incremental object detection (IOD) aims to extend detectors to new categories while retaining previously acquired knowledge. Existing methods often adopt a class incremental learni…

cs.CV2026

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

Zhipeng Xu, Zulong Chen, Qing Liu +6

Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-s…

cs.CV2026

Latent Visual Cache for Video Reasoning

Yongheng Zhang, Zhipeng Xu, Hao Wu +4

Video reasoning requires Large Multimodal Models (LMMs) to remain grounded in dense evidence, yet existing systems largely adopt "read-once, generate-many" paradigm, in which visua…