3 papers
cs.CV2026
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
Xingsong Ye, Yongkun Du, Jiaxin Zhang +5
WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially more challenging than general…
cs.CV2025
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
Yongyi Su, Haojie Zhang, Shijie Li +11
Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as genera…
cs.CV2024
PointSAM: Pointly-Supervised Segment Anything Model for Remote Sensing Images
Nanqing Liu, Xun Xu, Yongyi Su +2
Segment Anything Model (SAM) is an advanced foundational model for image segmentation, which is gradually being applied to remote sensing images (RSIs). Due to the domain gap betwe…