2 papers
cs.CV2026
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
Fangxu Yu, Ziyao Lu, Liqiang Niu +2
Grounding events in videos serves as a fundamental capability in video analysis. While Vision Language Models (VLMs) are increasingly employed for this task, existing approaches pr…
cs.CV2024
MaskMamba: A Hybrid Mamba-Transformer Model for Masked Image Generation
Wenchao Chen, Liqiang Niu, Ziyao Lu +2
Image generation models have encountered challenges related to scalability and quadratic complexity, primarily due to the reliance on Transformer-based backbones. In this study, we…