5 papers
UNICBench: UNIfied Counting Benchmark for MLLM
Chenggang Rong, Tao Han, Zhiyuan Zhao +5
Counting is a core capability for multimodal large language models (MLLMs), yet there is no unified counting dataset to rigorously evaluate this ability across image, text, and aud…
Real-Time Text Detection with Similar Mask in Traffic, Industrial, and Natural Scenes
Xu Han, Junyu Gao, Chuang Yang +2
Texts on the intelligent transportation scene include mass information. Fully harnessing this information is one of the critical drivers for advancing intelligent transportation. U…
Focus Entirety and Perceive Environment for Arbitrary-Shaped Text Detection
Xu Han, Junyu Gao, Chuang Yang +2
Due to the diversity of scene text in aspects such as font, color, shape, and size, accurately and efficiently detecting text is still a formidable challenge. Among the various det…
Spotlight Text Detector: Spotlight on Candidate Regions Like a Camera
Xu Han, Junyu Gao, Chuang Yang +2
The irregular contour representation is one of the tough challenges in scene text detection. Although segmentation-based methods have achieved significant progress with the help of…
Quantum-inspired Interpretable Deep Learning Architecture for Text Sentiment Analysis
Bingyu Li, Da Zhang, Zhiyuan Zhao +2
Text has become the predominant form of communication on social media, embedding a wealth of emotional nuances. Consequently, the extraction of emotional information from text is o…