Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models
Ting Chen, Geng Li, Guohao Chen +5
Contrastive decoding (CD) seeks to mitigate hallucinations in Large Vision-Language Models (LVLMs) by contrasting the output distributions of a standard model and a visually degrad…
cs.CV2025
Skeleton and Font Generation Network for Zero-shot Chinese Character Generation
Mobai Xue, Jun Du, Zhenrong Zhang +5
Automatic font generation remains a challenging research issue, primarily due to the vast number of Chinese characters, each with unique and intricate structures. Our investigation…
cs.CV2024
UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition
Zhenrong Zhang, Shuhang Liu, Pengfei Hu +4
In the digital era, table structure recognition technology is a critical tool for processing and analyzing large volumes of tabular data. Previous methods primarily focus on visual…