6 papers
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
Zhuchenyang Liu, Ziyi Wang, Yao Zhang +1
Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Prior compression routes either t…
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
Zhuchenyang Liu, Yao Zhang, Yu Xiao
Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a…
Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction
Yao Zhang, Zhuchenyang Liu, Yanlan He +2
Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences, enabling bidirectional search…
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
Zhuchenyang Liu, Yao Zhang, Yu Xiao
Vision-Language Model (VLM) based retrievers have advanced visual document retrieval (VDR) to impressive quality. They require the same multi-billion parameter encoder for both doc…
LingoMotion: An Interpretable and Unambiguous Symbolic Representation for Human Motion
Yao Zhang, Zhuchenyang Liu, Yu Xiao
Existing representations for human motion, such as MotionGPT, often operate as black-box latent vectors with limited interpretability and build on joint positions which can cause a…
Exploring Human-AI Collaboration in E-Textile Design: A Case Study on Flex Sensor Placement for Shoulder Motion Detection
Zhuchenyang Liu, Yao Zhang, Yalan He +4
Flex sensors are widely used in e-textiles for detecting joint motions and, subsequently, full-body movements. A critical initial step in utilizing these sensors is determining the…