6 papers · 1 filter
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
Zhuchenyang Liu, Yao Zhang, Yu Xiao
Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a…
Encoder-Free Human Motion Understanding via Structured Motion Descriptions
Yao Zhang, Zhuchenyang Liu, Thomas Ploetz +1
The world knowledge and reasoning capabilities of text-based large language models (LLMs) are advancing rapidly, yet current approaches to human motion understanding, including mot…
Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment
Zhuchenyang Liu, Yao Zhang, Yu Xiao
2D assembly diagrams are often abstract and hard to follow, creating a need for intelligent assistants that can monitor progress, detect errors, and provide step-by-step guidance.…
LingoMotion: An Interpretable and Unambiguous Symbolic Representation for Human Motion
Yao Zhang, Zhuchenyang Liu, Yu Xiao
Existing representations for human motion, such as MotionGPT, often operate as black-box latent vectors with limited interpretability and build on joint positions which can cause a…
Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction
Yao Zhang, Zhuchenyang Liu, Yanlan He +2
Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences, enabling bidirectional search…
Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval
Zhuchenyang Liu, Ziyu Hu, Yao Zhang +1
Recent Vision-Language Models (e.g., ColPali) enable fine-grained Visual Document Retrieval (VDR) but incur prohibitive multi-vector index storage overhead. Existing training-free…