9 papers
Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow
Jiangling Zhang, Shuxuan Gao, Zeyu Chen +2
Rapidly evolving Generative AI enables sophisticated visual text manipulations that increasingly evade current forensic detectors. Existing discriminative models often overfit spec…
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
Wenxuan Guo, Ziyuan Li, Meng Zhang +7
Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, existing VLA systems primarily re…
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
Zeyu Chen, Jie Li, Kai Han
Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross-modal information discrepancies and data scar…
Beyond Detection: A Structure-Aware Framework for Scene Text Tracking
Chenmin Yu, Liu Yu, Daiqing Wu +3
Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Although currently underexplored,…
DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding
Yichao Liu, Huawen Shen, Liu Yu +3
GUI agents powered by Multimodal Large Language Models (MLLMs) have demonstrated impressive capability in understanding and executing user instructions. However, accurately groundi…
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
Zeyu Chen, Fangmin Zhao, Yan Shu +3
Style-conditioned scene text generation faces unique challenges in extracting precise text styles from complex backgrounds and maintaining fine-grained style consistency across cha…