collaborators

10 papers

cs.IR2026

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation

Zhuchenyang Liu, Ziyi Wang, Yao Zhang +1

Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Prior compression routes either t…

cs.CV2026

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

Zhuchenyang Liu, Yao Zhang, Yu Xiao

Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a…

cs.CV2026

Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction

Yao Zhang, Zhuchenyang Liu, Yanlan He +2

Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences, enabling bidirectional search…

cs.LG2026

PJ-RoPE: A Fourier-Jet-Affine Position Space for Relative Attention

Yaobo Zhang

We organize relative-position mechanisms in attention as a learnable Fourier-Jet-Affine position space. The starting point is lag-shift dynamics: a relative-position kernel is a re…

cs.CV2026

Encoder-Free Human Motion Understanding via Structured Motion Descriptions

Yao Zhang, Zhuchenyang Liu, Thomas Ploetz +1

The world knowledge and reasoning capabilities of text-based large language models (LLMs) are advancing rapidly, yet current approaches to human motion understanding, including mot…

cs.IR2026

NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval

Zhuchenyang Liu, Yao Zhang, Yu Xiao

Vision-Language Model (VLM) based retrievers have advanced visual document retrieval (VDR) to impressive quality. They require the same multi-billion parameter encoder for both doc…