most citedARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.IR2026

From Saliency to Discriminability: Rank-Preserving Visual Token Pruning for VLM Rerankers

Siyi Liu, Hanjun Yang, Chenchen Zhang +7

Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning essential for practical deploymen…

cs.CV2026

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

Zirong Chen, Fuda Ye, Kuan Zhang +7

Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mo…

cs.IR2026

RePair: Turning Retrieval Failures into Counterfactual Hard Pairs

Siyi Liu, Xiaorong Zhu, Enjun Du +6

Vision-language retrieval with CLIP-style dual encoders achieves strong cross-modal performance, yet practical accuracy often hinges on localized semantic distinctions where top-ra…

cs.CV2026

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

Enjun Du, Siyi Liu, Zirong Chen +8

Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet exis…

cs.CV2026

Beyond Closed-Pool Video Retrieval: A Benchmark and Agent Framework for Real-World Video Search and Moment Localization

Tao Yu, Yujia Yang, Haopeng Jin +17

Traditional video retrieval benchmarks focus on matching precise descriptions to closed video pools, failing to reflect real-world searches characterized by fuzzy, multi-dimensiona…

cs.CV20251 cited

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li +15

Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal…