19 citations · 77 across the 31 of their papers we have counts for
6 papers · 1 filter
Magnitude-Direction Decoupling for Fast Video Generation with Flow Matching Models
Haonan Xu, Feiyang Chen, Songkui Chen +5
Flow matching models for video generation achieve impressive performance but suffer from high computational overhead due to iterative denoising. In fact, the original model is not…
MLLM-Guided Semantic Correction for Text-to-Video Generation
Junhao Chen, Zheqi Lv, Keting Yin +6
Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic err…
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
Yunxin Li, Zhenyu Liu, Zitao Li +19
Reasoning lies at the heart of intelligence, shaping the ability to make decisions, draw conclusions, and generalize across domains. In artificial intelligence, as systems increasi…
Recognizing Unseen Objects via Multimodal Intensive Knowledge Graph Propagation
Likang Wu, Zhi Li, Hongke Zhao +5
Zero-Shot Learning (ZSL), which aims at automatically recognizing unseen objects, is a promising learning paradigm to understand new real-world knowledge for machines continuously.…
VLAD-VSA: Cross-Domain Face Presentation Attack Detection with Vocabulary Separation and Adaptation
Jiong Wang, Zhou Zhao, Weike Jin +5
For face presentation attack detection (PAD), most of the spoofing cues are subtle, local image patterns (e.g., local image distortion, 3D mask edge and cut photo edges). The repre…
Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding
Zhu Zhang, Zhou Zhao, Zhijie Lin +2
Spatio-temporal video grounding aims to retrieve the spatio-temporal tube of a queried object according to the given sentence. Currently, most existing grounding methods are restri…