1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CV2025
AVGGT: Rethinking Global Attention for Accelerating VGGT
Xianbing Sun, Zhikai Zhu, Zhengyu Lou +5
Models such as VGGT and have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-att…
cs.CL2024
UNER: A Unified Prediction Head for Named Entity Recognition in Visually-rich Documents
Yi Tu, Chong Zhang, Ya Guo +4
The recognition of named entities in visually-rich documents (VrD-NER) plays a critical role in various real-world scenarios and applications. However, the research in VrD-NER face…
cs.CL2023★ 1 cited
Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction
Chong Zhang, Ya Guo, Yi Tu +5
Recent advances in multimodal pre-trained models have significantly improved information extraction from visually-rich documents (VrDs), in which named entity recognition (NER) is…