activity
20182025
most citedCATN: Cross-Domain Recommendation for Cold-Start Users via Aspect Transfer Network

194 citations · 403 across the 49 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework

Jing Wang, Fengzhuo Zhang, Xiaoli Li +5

Auto-Regressive Video Diffusion Models (AR-VDMs) have shown strong capabilities in generating long, photorealistic videos, but suffer from two key limitations: (i) history forgetti…

cs.CV2024

MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Yubo Ma, Yuhang Zang, Liangyu Chen +13

Understanding documents with rich layouts and multi-modal components is a long-standing and practical task. Recent Large Vision-Language Models (LVLMs) have made remarkable strides…

cs.CV2024

Multi-scale 2D Temporal Map Diffusion Models for Natural Language Video Localization

Chongzhi Zhang, Mingyuan Zhang, Zhiyang Teng +5

Natural Language Video Localization (NLVL), grounding phrases from natural language descriptions to corresponding video segments, is a complex yet critical task in video understand…

cs.CV2023

MS-DETR: Natural Language Video Localization with Sampling Moment-Moment Interaction

Jing Wang, Aixin Sun, Hao Zhang +1

Given a query, the task of Natural Language Video Localization (NLVL) is to localize a temporal moment in an untrimmed video that semantically matches the query. In this paper, we…

cs.CV20216 cited

Towards Debiasing Temporal Sentence Grounding in Video

Hao Zhang, Aixin Sun, Wei Jing +1

The temporal sentence grounding in video (TSGV) task is to locate a temporal moment from an untrimmed video, to match a language query, i.e., a sentence. Without considering bias i…