7 papers
Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores
Zeyu Zhang, Bradly C. Stadie
The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff. We show this check is uninformative. Four flagship models fail…
BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation
Zeyu Zhang, Jinyuan Mao, Shuning Chang +5
Long video generation is a critical step toward building realistic world models, requiring both high visual fidelity and long-range interaction consistency. Recent autoregressive d…
Towards Error-Free Long Video Generation
Shuning Chang, Weihua Chen, Jiasheng Tang +8
Recent advances in video generation have made minute-level synthesis possible; however, generating long videos remains challenging due to error accumulation, attribute drift, and t…
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
Akide Liu, Jinbo Xing, Chaojie Mao +8
Minute-scale cinematic video generation is a central challenge for generative video models. Existing paradigms address only fragments of this challenge: single-shot extrapolation p…
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction
Weijie Wang, Zimu Li, Jinchuan Shi +5
Sparse-view 3D reconstruction is increasingly addressed with feed-forward splatting networks that predict explicit primitives directly from images. Yet most existing methods remain…
All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting
Zeyu Zhang, Ryan Chen, Bradly C. Stadie
Backtesting LLMs on resolved events assumes models reason only from pre-cutoff knowledge, yet pretrained models inevitably leak post-cutoff knowledge. We introduce a claim-level ev…