2 papers
cs.CV2026
One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting
Rui Tang, Wentao Yang, Peirong Zhang +4
Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as a promising formulation for…
cs.CV2026
SemanticFace: Semantic Facial Action Estimation via Semantic Distillation in Interpretable Space
Zejian Kang, Kai Zheng, Yuanchen Fei +3
Facial action estimation from a single image is often formulated as predicting or fitting parameters in compact expression spaces, which lack explicit semantic interpretability. Ho…