9 citations · 17 across the 3 of their papers we have counts for
4 papers · 1 filter
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
Wenbin An, Jiahao Nie, Yaqiang Wu +3
By integrating the perception capabilities of multimodal encoders with the generative power of Large Language Models (LLMs), Multimodal Large Language Models (MLLMs), exemplified b…
Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution
Jiapeng Wang, Chongyu Liu, Lianwen Jin +6
Visual information extraction (VIE) has attracted considerable attention recently owing to its various advanced applications such as document understanding, automatic marking and i…
Decoupled Attention Network for Text Recognition
Tianwei Wang, Yuanzhi Zhu, Lianwen Jin +5
Text recognition has attracted considerable research interests because of its various applications. The cutting-edge text recognition methods are based on attention mechanisms. How…
Omnidirectional Scene Text Detection with Sequential-free Box Discretization
Yuliang Liu, Sheng Zhang, Lianwen Jin +3
Scene text in the wild is commonly presented with high variant characteristics. Using quadrilateral bounding box to localize the text instance is nearly indispensable for detection…