4 citations · 7 across the 13 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
How to Utilize Complementary Vision-Text Information for 2D Structure Understanding
Jiancheng Dong, Pengyue Jia, Derong Xu +9
LLMs typically linearize 2D tables into 1D sequences to fit their autoregressive architecture, which weakens row-column adjacency and other layout cues. In contrast, purely visual…
cs.CV2024
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
Qian Li, Lixin Su, Jiashu Zhao +6
Text-video retrieval is a challenging task that aims to identify relevant videos given textual queries. Compared to conventional textual retrieval, the main obstacle for text-video…