4 papers
TempCloze: Can Video-LLMs Identify the Missing Middle?
Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu +4
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To…
SlotMem: Character-Addressable Internal Memory for Narrative Long Video Generation
Yilai Liu, Xin Zhang, Shiyuan Zhang +1
Maintaining recurring character identities across scene transitions and long temporal gaps is a central challenge in narrative long video generation. Methods targeting global consi…
MIDiff: Tackling Sparsity and Imbalance in Mobile Usage Generation via Multivariate-Imaging Diffusion
Yilai Liu, Shiyuan Zhang, Hongyang Du
Mobile usage traces are critical for tasks such as user behavior prediction and app recommendation, yet their use is constrained by privacy restrictions and costly large-scale data…
U-MASK: User-adaptive Spatio-Temporal Masking for Personalized Mobile AI Applications
Shiyuan Zhang, Yilai Liu, Yuwei Du +3
Personalized mobile artificial intelligence applications are widely deployed, yet they are expected to infer user behavior from sparse and irregular histories under a continuously…