3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CV2025
TextVidBench: A Benchmark for Long Video Scene Text Understanding
Yangyang Zhong, Ji Qi, Yuan Yao +5
Despite recent progress on the short-video Text-Visual Question Answering (ViteVQA) task - largely driven by benchmarks such as M4-ViteVQA - existing datasets still suffer from lim…
cs.CV2025
Adept: Annotation-Denoising Auxiliary Tasks with Discrete Cosine Transform Map and Keypoint for Human-Centric Pretraining
Weizhen He, Yunfeng Yan, Shixiang Tang +4
Human-centric perception is the core of diverse computer vision tasks and has been a long-standing research focus. However, previous research studied these human-centric tasks indi…
cs.CL2024★ 3 cited
Unveiling the Impact of Multi-Modal Interactions on User Engagement: A Comprehensive Evaluation in AI-driven Conversations
Lichao Zhang, Jia Yu, Shuai Zhang +10
Large Language Models (LLMs) have significantly advanced user-bot interactions, enabling more complex and coherent dialogues. However, the prevalent text-only modality might not fu…